Skip to content

InferenceSambaNova Systems

SambaNova Cloud

Fast inference for open models on custom RDU chips.

Category
Inference
Pricing
FREEMIUM
Hosting
Cloud
Platforms
WebAPI
Models
Multi-model
Verified
Jun 7, 2026

Inference cloud running open-weight models — Llama, DeepSeek, Qwen, gpt-oss — on SambaNova's RDU hardware at hundreds of tokens per second, including full-precision Llama 405B. Provides an OpenAI-compatible API with a free tier and pay-per-token pricing.

Capabilities 2

What it actually does — grouped by capability family.

  • Model inference / serving (primary capability)
  • Multi-model access (secondary capability)

Pros & cons

  • Serves Llama, DeepSeek, Qwen, gpt-oss
  • Hundreds of tokens/sec on RDU chips
  • OpenAI-compatible API
  • Free tier to start
  • Open-weight catalog only
  • No fine-tuning/custom hosting like GPU clouds
  • Smaller model selection than rivals

Tags

View all Inference
  • View Groq details
    InferenceFREEMIUM

    Groq

    Groq

    Low-latency inference for open-weights models on custom LPU chips.

    GroqCloud serves open-weights models (Llama, DeepSeek, Qwen, Kimi) on Groq's purpose-built LPU hardware, hitting hundreds of tokens per second where GPUs manage tens. OpenAI-compatible API with a free tier; the default when token latency is the product.

    Hundreds of tokens/sec on open models
    Curated open-weight models only
    • inference
    • low-latency
    • lpu
    • open-weights
  • View Cerebras details
    InferenceFREEMIUM

    Cerebras

    Cerebras Systems

    Wafer-scale inference cloud for open models.

    Inference cloud that serves open-weight models such as Llama, Qwen, DeepSeek, and gpt-oss on Cerebras's wafer-scale CS-3 hardware, reaching token throughput far above GPU clouds. Exposes an OpenAI-compatible API with a free daily tier and pay-per-token pricing.

    Highest tokens/sec in the market
    Smaller model catalog than Groq/Together
    • inference
    • fast-inference
    • wafer-scale
    • open-models
  • View Together AI details
    InferenceFREEMIUM

    Together AI

    Together

    Hosted inference and fine-tuning for open-weights models.

    Hosted inference and fine-tuning across hundreds of open-weights models (Llama, Mistral, DeepSeek, Qwen, etc.). Strong pricing for inference-at-scale; LoRA + full fine-tuning supported.

    LoRA and full fine-tuning
    Open models only, no frontier closed models
    • inference
    • fine-tuning
    • open-weights
    • lora
  • View Fireworks AI details
    InferenceFREEMIUM

    Fireworks AI

    Fireworks AI

    Fast inference + fine-tuning. Production deployments at scale.

    Optimized inference platform for open-weights models with strong latency numbers and serverless + dedicated deployment options. Fine-tuning supported; vision and audio models alongside text.

    Custom FireAttention inference stack
    Usage pricing scales with traffic
    • inference
    • fine-tuning
    • low-latency
    • production