Skip to content

VisionM87 Labs

Moondream

Tiny open vision-language model for efficient image understanding.

Category
Vision
Pricing
FREEMIUM
Source
Open core
Hosting
Hybrid
Platforms
WebAPI
Models
Self-contained (on-device)
Verified
Jun 7, 2026

An open-weights family of small vision-language models for captioning, visual Q&A, pointing, counting, and object detection — small enough to run on-device (checkpoints down to 0.5B on Hugging Face). Run it locally with the Photon engine, or call Moondream Cloud's OpenAI-compatible API with a free monthly credit tier and pay-per-image pricing.

Capabilities 5

What it actually does — grouped by capability family.

  • Model inference / serving (secondary capability)
  • Fine-tuning / training (secondary capability)
  • OCR / scanned-document extraction (secondary capability)
  • Object detection (primary capability)
  • Image classification (primary capability)

Pros & cons

  • Open-weights, free to self-host
  • Runs on-device with Photon engine
  • Does pointing, counting, detection
  • OpenAI-compatible cloud API option
  • Small models trail frontier VLMs on hard tasks
  • Narrower than large multimodal LLMs
  • Cloud tier is pay-per-image

Tags

View all Vision
  • View Roboflow details
    VisionFREEMIUM

    Roboflow

    Roboflow

    Vision MLOps end-to-end. Annotate, train, deploy.

    Annotation tooling, auto-labelling, hosted training, and edge deployment for computer-vision projects. Strong default when you're shipping a custom vision model rather than reaching for a multimodal LLM.

    End-to-end vision MLOps
    Free tier caps usage and privacy
    • annotation
    • training
    • deployment
    • edge
  • View LandingAI details
    VisionFREEMIUM

    LandingAI

    LandingAI

    Build vision detectors and agents from a few labeled examples.

    Build vision applications with a labelling-light workflow — point at examples, get a deployable detector. Recently extended into vision agents that reason over images and PDFs without bespoke training.

    Fast path to a deployable detector
    Less control than custom model training
    • visual-prompting
    • agents
    • document-ai
    • no-code
  • View TwelveLabs details
    VisionFREEMIUM

    TwelveLabs

    TwelveLabs

    Video intelligence API: search, classify, and summarize video.

    Video understanding platform built on its own multimodal foundation models — Marengo for embeddings and semantic search, Pegasus for generative tasks like summaries and captions. Developers index video once and run natural-language search, classification, and analysis via API. Free tier with usage-based pricing beyond it.

    Marengo embeddings + Pegasus generation
    Proprietary, closed models
    • video-understanding
    • search
    • multimodal
    • embeddings
    • +1