Skip to content

InferenceLM Studio

LM Studio

Desktop app to discover, download, and run local LLMs privately.

Category
Inference
Pricing
FREE
Hosting
Local
Platforms
macOSWindowsLinuxCLIAPI
Models
Multi-model
Verified
Jun 6, 2026

A GUI for running open-weight models on your own hardware — browse and download GGUF/MLX models, chat offline, and expose an OpenAI- and Anthropic-compatible local server for your apps. Includes RAG over local files, MCP tool-use support, and dual llama.cpp + Apple MLX runtimes. Free for personal and commercial use; the app itself is proprietary.

Capabilities 4

What it actually does — grouped by capability family.

  • Model inference / serving (primary capability)
  • Chat with documents (secondary capability)
  • RAG pipeline (secondary capability)
  • Embeddings (secondary capability)

Pros & cons

  • Polished desktop GUI
  • In-app Hugging Face model search
  • RAG over local files + MCP tool-use
  • Free for personal + commercial use
  • App itself is closed source
  • Heavier (Electron) than Ollama
  • Slower model loads vs Ollama

Tags

View all Inference
  • View Ollama details
    InferenceFREEMIUMOpen core

    Ollama

    Ollama

    Run open-weight LLMs locally with one command. OpenAI-compatible API.

    The de-facto way to pull and run open-weight models (Llama, Qwen, Gemma, DeepSeek, gpt-oss) on your own machine — no API key, no data leaving the device. Ships native macOS/Windows/Linux apps, an OpenAI-compatible server, and official Python/JS libraries. MIT-licensed and free locally; an optional paid Ollama Cloud runs larger models.

    One-command pull-and-run
    Local performance bound by your hardware
    • local
    • open-source
    • llm-runner
    • self-hosted
  • View Jan details
    AssistantFREEOSS

    Jan

    Menlo Research

    ChatGPT alternative that runs 100% offline on your computer.

    An open-source desktop assistant that runs AI models 100% offline on your own machine, bundling a local model runner so you can download and chat with open models like Llama, Gemma, and Qwen. It also supports bring-your-own keys for cloud providers when you want them, and exposes a local OpenAI-compatible server. Fully open source under Apache 2.0 with no paid tier.

    Local-first; data stays on your machine
    More setup friction than LM Studio
    • offline
    • local
    • open-source
    • privacy
    • +1
  • View AnythingLLM details
    AssistantFREEMIUMOpen core

    AnythingLLM

    Mintplex Labs

    All-in-one private AI app for chatting with your documents, with agents.

    An all-in-one application for private, ChatGPT-style chat over your own documents, with built-in RAG, AI agents, and multi-user workspaces. Runs as a local desktop app (Mac/Windows/Linux) or self-hosted via Docker, and supports 40+ LLM providers plus local models with your own keys. Open source under MIT; Mintplex Labs also offers a paid hosted instance, making it freemium.

    MIT open source
    RAG quality depends on your setup
    • rag
    • documents
    • self-hosted
    • open-source
    • +1
  • View vLLM details
    InferenceFREEOSS

    vLLM

    vLLM Project

    High-throughput, memory-efficient inference engine for LLMs.

    A serving engine for large language and vision-language models, originally from UC Berkeley's Sky Computing Lab. Its PagedAttention KV-cache management and continuous batching deliver high throughput on commodity GPUs. Now a community project with 1000s of contributors and an OpenAI-compatible server.

    Serves most Hugging Face transformer models
    You manage the GPU infrastructure
    • inference
    • model-serving
    • gpu
    • open-source
    • +1