Skip to content

EvalGiskard

Giskard

Open-source evaluation and red-teaming for LLM agents and RAG apps.

Categories
EvalSecurity
Pricing
FREEMIUM
Source
Open core
Hosting
Hybrid
Platforms
WebAPI
Models
Model-agnostic
Verified
Jun 8, 2026

Giskard is an open-source (Apache-2.0) Python library for testing LLMs, RAG pipelines, and ML models — its Scan automatically surfaces hallucinations, prompt injection, bias, and other vulnerabilities, while red-teaming agents run multi-turn adversarial attacks across dozens of probes. The paid Giskard Hub adds team collaboration, continuous testing, and scheduled scans. The team also publishes the open Phare LLM safety benchmark.

Capabilities 3

What it actually does — grouped by capability family.

  • Red-teaming (primary capability)
  • AI security scanning (primary capability)
  • LLM evaluation (secondary capability)

Pros & cons

  • Automatic vulnerability scan
  • Multi-turn red-teaming agents
  • Covers LLMs, RAG apps, and ML models
  • Publishes the open Phare safety benchmark
  • Python-library learning curve
  • Collaboration features are paid (Hub)
  • Less focused on production tracing

Tags

View all Eval
  • View Promptfoo details
    EvalFREEOSS

    Promptfoo

    Promptfoo

    LLM eval CLI with rubric scoring and golden sets.

    YAML-driven eval harness. Pair a prompt with a goldset, define rubrics, run across multiple models in CI. Strong for catching prompt regressions before they hit production.

    YAML-driven, version-controllable evals
    CLI-first, less of a hosted UI
    • eval
    • ci
    • rubric
    • open-source
  • View DeepEval details
    EvalFREEMIUMOpen core

    DeepEval

    Confident AI

    Pytest-style framework for evaluating LLM apps in CI.

    Open-source (Apache 2.0) framework for evaluating LLM apps the way Pytest tests code — assertions backed by 50+ ready metrics spanning LLM-as-judge, RAG, agents, conversation, and safety. Plugs into LangChain, CrewAI, OpenAI Agents and more. Confident AI is the paid cloud platform that adds test management, dashboards, and observability on top.

    Assertions run in your CI pipeline
    LLM-as-judge adds cost
    • eval
    • open-source
    • llm-as-judge
    • rag
    • +1
  • View Patronus AI details
    EvalFREEMIUM

    Patronus AI

    Patronus AI

    Automated evaluation, guardrails, and monitoring for AI systems.

    Platform for evaluating, guarding, and monitoring LLM and agent applications across the deployment lifecycle. Anchored by research-backed evaluator models — Lynx (hallucination detection), GLIDER (LLM judge), and Percival (agent-trace debugger). Offers a self-serve API with free credits, usage-based pricing, and enterprise plans.

    Research-backed Lynx, GLIDER, and Percival models
    Cloud-only; no self-host
    • eval
    • guardrails
    • monitoring
    • hallucination
    • +1
  • View Lakera details
    SecurityFREEMIUM

    Lakera

    Lakera (Check Point)

    Real-time guardrails against prompt injection and jailbreaks for AI apps.

    Lakera Guard sits between users and LLMs as a low-latency security layer, detecting and blocking direct and indirect prompt injection, jailbreaks, and system-prompt extraction across 100+ languages. Its models are trained on adversarial data from Gandalf, Lakera's prompt-injection game. Acquired by Check Point in 2025.

    Detection sharpened by the Gandalf game
    Behavioral detection risks false positives
    • prompt-injection
    • guardrails
    • llm-security
    • jailbreak