Skip to content

Data OpsTreeverse

lakeFS

Git-like version control for data lakes over your existing object storage.

Category
Data Ops
Pricing
FREEMIUM
Source
Open core
Hosting
Hybrid
Platforms
WebCLIAPI
Models
Model-agnostic
Verified
Jun 21, 2026

Open-source data version control that turns object storage (S3, GCS, Azure Blob, MinIO) into Git-like repositories. Teams branch, commit, merge, and roll back petabyte-scale data lakes for isolated experimentation, reproducible ML pipelines, data-quality gates, and compliance lineage — without copying data. Integrates with Spark, Trino, Databricks, Delta Lake, and Iceberg.

Capabilities 1

What it actually does — grouped by capability family.

  • ETL / data pipeline (primary capability)

Pros & cons

  • Open source (Apache 2.0)
  • Isolated experiments and reproducible pipelines
  • Rollback and data-quality gates
  • Integrates with Spark, Trino, Iceberg, Delta
  • Managed Cloud and self-host options
  • Operational overhead to self-host
  • Aimed at data-lake-scale teams
  • Advanced features gated to paid tiers

Tags

Further reading

View all Data Ops
  • View DVC details
    Data OpsFREEOSS

    DVC

    lakeFS

    Git extension for versioning data, models, and ML experiments.

    DVC (Data Version Control) brings software-engineering practices to machine learning: it versions datasets, models, and pipelines alongside code in any Git repository, storing large files in your own remote storage while keeping lightweight pointers in Git. Enables reproducible experiments, data/model lineage, and pipeline orchestration from the command line.

    Free and open source
    CLI-centric learning curve
    • data-versioning
    • mlops
    • reproducibility
    • ml-pipelines
    • +2
  • View Deep Lake details
    Vector DBFREEMIUMOpen core

    Deep Lake

    Activeloop

    Multimodal database for AI — vectors plus raw data, versioned.

    Deep Lake, by Activeloop, is a database for AI that stores vectors alongside raw multimodal data — text, image, video, audio, and metadata — in a single version-controlled format. It supports cross-modal queries and can stream data straight into training, and its open-source core can be self-hosted or run as a managed service. The newer Deep Lake PG unifies a serverless Postgres with its vector and tensor engine.

    Open-source core (self-host or cloud)
    Smaller community than Pinecone/Qdrant
    • vector-database
    • multimodal
    • data-lake
    • versioning
  • View MLflow details
    ObservabilityFREEOSS

    MLflow

    Linux Foundation

    Open-source platform for the ML and GenAI lifecycle.

    MLflow is an open-source platform for managing the full machine-learning and GenAI lifecycle — experiment tracking, model registry, deployment, and, more recently, LLM/agent observability. Its GenAI stack adds OpenTelemetry-based tracing, systematic evaluation with built-in metrics and LLM judges, and prompt versioning. Framework- and provider-agnostic, it runs on your own infrastructure with no vendor lock-in.

    Fully open source, no lock-in
    Self-hosting adds operational overhead
    • llmops
    • tracing
    • evaluation
    • mlops
    • +1