Skip to content

InferenceHyperbolic

Hyperbolic

Open-access AI cloud: serverless inference + a GPU marketplace.

Categories
InferenceInfra
Pricing
FREEMIUM
Hosting
Cloud
Platforms
APIWeb
Models
Multi-model
Verified
Jun 9, 2026

An AI cloud offering serverless inference for open models (Llama, Qwen, DeepSeek, SDXL, Flux) behind an OpenAI-compatible API, alongside an on-demand GPU marketplace for H100/H200 rentals. It aggregates idle and reserved compute to price inference and GPU hours below centralized clouds. Aimed at developers training, fine-tuning, and serving open-weights models.

Capabilities 4

What it actually does — grouped by capability family.

  • Model inference / serving (primary capability)
  • Multi-model access (primary capability)
  • GPU compute (secondary capability)
  • LLM gateway / routing (secondary capability)

Pros & cons

  • Serverless inference + GPU marketplace
  • On-demand H100/H200 GPU rentals
  • OpenAI-compatible API
  • Open models: Llama, Qwen, DeepSeek, FLUX
  • Marketplace supply reliability varies
  • Open-weights only, no frontier closed models
  • Smaller/newer than AWS-scale clouds
  • Less enterprise tooling

Tags

View all Inference
  • View Together AI details
    InferenceFREEMIUM

    Together AI

    Together

    Hosted inference and fine-tuning for open-weights models.

    Hosted inference and fine-tuning across hundreds of open-weights models (Llama, Mistral, DeepSeek, Qwen, etc.). Strong pricing for inference-at-scale; LoRA + full fine-tuning supported.

    LoRA and full fine-tuning
    Open models only, no frontier closed models
    • inference
    • fine-tuning
    • open-weights
    • lora
  • View Fireworks AI details
    InferenceFREEMIUM

    Fireworks AI

    Fireworks AI

    Fast inference + fine-tuning. Production deployments at scale.

    Optimized inference platform for open-weights models with strong latency numbers and serverless + dedicated deployment options. Fine-tuning supported; vision and audio models alongside text.

    Custom FireAttention inference stack
    Usage pricing scales with traffic
    • inference
    • fine-tuning
    • low-latency
    • production
  • View Runpod details
    InferencePAID

    Runpod

    Runpod

    GPU cloud for AI — on-demand instances and serverless inference.

    Runpod is an AI developer cloud for renting GPUs on demand or running auto-scaling serverless inference endpoints. Serverless workers bill by the millisecond, scale to zero when idle, and advertise sub-200ms cold starts; on-demand Pods and multi-node Clusters cover training and long-running jobs. A Community Cloud tier offers cheaper, peer-sourced GPUs alongside the vendor-operated Secure Cloud.

    Serverless auto-scaling inference
    Community Cloud less reliable/secure
    • gpu-cloud
    • serverless
    • inference
    • deployment
    • +1
  • View Replicate details
    InferenceFREEMIUM

    Replicate

    Replicate

    Run, fine-tune, and deploy thousands of open models via one API.

    A platform to run open-source models with one API call — image, video, audio, and language — plus fine-tuning and custom deploys with pay-per-second billing. No infra to manage.

    Image, video, audio, and language models
    Cold starts on less-popular models
    • model-hosting
    • fine-tuning
    • api
    • open-source