Skip to content

InferenceCerebras Systems

Cerebras

Wafer-scale inference cloud for open models.

Category
Inference
Pricing
FREEMIUM
Hosting
Cloud
Platforms
WebAPI
Models
Multi-model
Verified
Jun 7, 2026

Inference cloud that serves open-weight models such as Llama, Qwen, DeepSeek, and gpt-oss on Cerebras's wafer-scale CS-3 hardware, reaching token throughput far above GPU clouds. Exposes an OpenAI-compatible API with a free daily tier and pay-per-token pricing.

Capabilities 3

What it actually does — grouped by capability family.

  • Model inference / serving (primary capability)
  • Multi-model access (secondary capability)
  • Fine-tuning / training (secondary capability)

Pros & cons

  • Highest tokens/sec in the market
  • Low time-to-first-token (~80-150ms)
  • 2-3x faster end-to-end in agent loops
  • OpenAI-compatible API, free daily tier
  • Smaller model catalog than Groq/Together
  • Less mature ecosystem and client libs
  • Occasional capacity limits under demand

Tags

View all Inference
  • View Groq details
    InferenceFREEMIUM

    Groq

    Groq

    Low-latency inference for open-weights models on custom LPU chips.

    GroqCloud serves open-weights models (Llama, DeepSeek, Qwen, Kimi) on Groq's purpose-built LPU hardware, hitting hundreds of tokens per second where GPUs manage tens. OpenAI-compatible API with a free tier; the default when token latency is the product.

    Hundreds of tokens/sec on open models
    Curated open-weight models only
    • inference
    • low-latency
    • lpu
    • open-weights
  • View SambaNova Cloud details
    InferenceFREEMIUM

    SambaNova Cloud

    SambaNova Systems

    Fast inference for open models on custom RDU chips.

    Inference cloud running open-weight models — Llama, DeepSeek, Qwen, gpt-oss — on SambaNova's RDU hardware at hundreds of tokens per second, including full-precision Llama 405B. Provides an OpenAI-compatible API with a free tier and pay-per-token pricing.

    Serves Llama, DeepSeek, Qwen, gpt-oss
    Open-weight catalog only
    • inference
    • fast-inference
    • open-models
    • rdu
  • View Together AI details
    InferenceFREEMIUM

    Together AI

    Together

    Hosted inference and fine-tuning for open-weights models.

    Hosted inference and fine-tuning across hundreds of open-weights models (Llama, Mistral, DeepSeek, Qwen, etc.). Strong pricing for inference-at-scale; LoRA + full fine-tuning supported.

    LoRA and full fine-tuning
    Open models only, no frontier closed models
    • inference
    • fine-tuning
    • open-weights
    • lora
  • View Fireworks AI details
    InferenceFREEMIUM

    Fireworks AI

    Fireworks AI

    Fast inference + fine-tuning. Production deployments at scale.

    Optimized inference platform for open-weights models with strong latency numbers and serverless + dedicated deployment options. Fine-tuning supported; vision and audio models alongside text.

    Custom FireAttention inference stack
    Usage pricing scales with traffic
    • inference
    • fine-tuning
    • low-latency
    • production