Skip to content

InferenceGroq

Groq

Low-latency inference for open-weights models on custom LPU chips.

Categories
InferenceInfra
Pricing
FREEMIUM
Hosting
Cloud
Platforms
APIWeb
Models
Multi-model
Verified
Jun 6, 2026

GroqCloud serves open-weights models (Llama, DeepSeek, Qwen, Kimi) on Groq's purpose-built LPU hardware, hitting hundreds of tokens per second where GPUs manage tens. OpenAI-compatible API with a free tier; the default when token latency is the product.

Capabilities 4

What it actually does — grouped by capability family.

  • Model inference / serving (primary capability)
  • Multi-model access (primary capability)
  • Transcription (STT) (secondary capability)
  • Speech synthesis (TTS) (secondary capability)

Pros & cons

  • Hundreds of tokens/sec on open models
  • Sub-100ms time-to-first-token
  • Deterministic, low-variance latency
  • OpenAI-compatible API with free tier
  • Curated open-weight models only
  • No frontier closed models (GPT/Claude)
  • SRAM limits large context windows
  • Rate limits during peak demand

Tags

View all Inference
  • View Cerebras details
    InferenceFREEMIUM

    Cerebras

    Cerebras Systems

    Wafer-scale inference cloud for open models.

    Inference cloud that serves open-weight models such as Llama, Qwen, DeepSeek, and gpt-oss on Cerebras's wafer-scale CS-3 hardware, reaching token throughput far above GPU clouds. Exposes an OpenAI-compatible API with a free daily tier and pay-per-token pricing.

    Highest tokens/sec in the market
    Smaller model catalog than Groq/Together
    • inference
    • fast-inference
    • wafer-scale
    • open-models
  • View SambaNova Cloud details
    InferenceFREEMIUM

    SambaNova Cloud

    SambaNova Systems

    Fast inference for open models on custom RDU chips.

    Inference cloud running open-weight models — Llama, DeepSeek, Qwen, gpt-oss — on SambaNova's RDU hardware at hundreds of tokens per second, including full-precision Llama 405B. Provides an OpenAI-compatible API with a free tier and pay-per-token pricing.

    Serves Llama, DeepSeek, Qwen, gpt-oss
    Open-weight catalog only
    • inference
    • fast-inference
    • open-models
    • rdu
  • View Fireworks AI details
    InferenceFREEMIUM

    Fireworks AI

    Fireworks AI

    Fast inference + fine-tuning. Production deployments at scale.

    Optimized inference platform for open-weights models with strong latency numbers and serverless + dedicated deployment options. Fine-tuning supported; vision and audio models alongside text.

    Custom FireAttention inference stack
    Usage pricing scales with traffic
    • inference
    • fine-tuning
    • low-latency
    • production
  • View Together AI details
    InferenceFREEMIUM

    Together AI

    Together

    Hosted inference and fine-tuning for open-weights models.

    Hosted inference and fine-tuning across hundreds of open-weights models (Llama, Mistral, DeepSeek, Qwen, etc.). Strong pricing for inference-at-scale; LoRA + full fine-tuning supported.

    LoRA and full fine-tuning
    Open models only, no frontier closed models
    • inference
    • fine-tuning
    • open-weights
    • lora