Apps tagged with 'LLM Evaluation'

All apps in Apps tagged with 'LLM Evaluation' category. Use the filters below to narrow down your search. 
Copy a direct link to this comment to your clipboard
  1. Langfuse icon
     1 like

    Open source LLM engineering platform: LLM Observability, metrics, evals, prompt management, playground, datasets. Integrates with OpenTelemetry, Langchain, OpenAI SDK, LiteLLM, and more.

    Cost / License

    • Freemium
    • Proprietary

    Platforms

    • Online
    • Self-Hosted
    • Software as a Service (SaaS)
    • Docker
    • Cloudron
    • DE flagGermany
    • European Union flagEU
    Create manual annotations to provide feedback, corrections, and improvements to your LLM outputs. Use annotations to build high-quality datasets and set a baseline for automated evals.
    Run online/offline evals, via UI (experiment with prompts/models) and via SDKs (experiment with end-to-end application). Build datasets from traces to continuously improve your evals. View results in UI.
    Experiment with different prompts, models, and parameters in an interactive playground. Compare outputs, iterate on prompts, and save successful configurations to prompt management.
    +3
    Version-control prompts collaboratively, deploy/roll-back instantly to different environments, support for templates, variables, and A/B testing. Cached client-side for 0 latency/availability impact.
    22 alternatives
  2. Zespan icon
     1 like

    The reliability platform for AI agents - trace, evaluate, guard, and control the cost of every LLM call and agent handoff in production.

    Cost / License

    • Freemium
    • Proprietary

    Platforms

    • Software as a Service (SaaS)
    • Self-Hosted
    • Online
    • Docker
    • IN flagIndia
    Zespan screenshot 1
    Zespan screenshot 1
    Zespan screenshot 2
    +11
    Zespan screenshot 3
    4 alternatives
  3. AIQ-X icon
     Like

    Test AI models yourself, privately, with a standardized benchmark, and get both technical scores AND practical recommendations.

    Cost / License

    • Free
    • Proprietary

    Platforms

    • Online
    • US flagUnited States
    Actionable Insights & Recommendations
  4. LightEval icon
     Like

    LightEval is a lightweight LLM evaluation suite that Hugging Face has been using internally with the recently released LLM data processing library datatrove and LLM training library nanotron.

    Cost / License

    • Free
    • Open Source (MIT)

    Platforms

    • Self-Hosted
    • Python
    3 alternatives
  5. Lisapet.ai is the next-level AI product development platform that empowers teams to prototype, test, and ship robust AI features 10x faster.

    Cost / License

    • Paid
    • Proprietary

    Platforms

    • Online
    Lisapet.ai Thumbnail
    Lisapet.ai Demo
    1 alternatives
  6. Promptfoo icon
     Like

    Open-source tool for automated LLM prompt, agent, and RAG evaluation, supporting red teaming, multi-model comparisons, CI/CD, CLI, and vulnerability scanning.

    Cost / License

    • Freemium
    • Open Source (MIT)

    Platforms

    • Online
    • Self-Hosted
    • US flagUnited States
    Promptfoo screenshot 1
    Promptfoo screenshot 1
    Promptfoo screenshot 2
    +1
    Promptfoo screenshot 3
    6 alternatives
  7.  Like

    Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.

    Cost / License

    • Freemium
    • Source Available

    Platforms

    • Mac
    • Windows
    • Linux
    • Python
    • CA flagCanada
    Kiln AI screenshot 1
    8 alternatives
  8. Respan AI icon
     Like

    Unified platform for routing LLM traffic, observability, evals, prompt management, tracing, and spend controls, reducing dashboard switching time.

    Cost / License

    • Free
    • Proprietary

    Platforms

    • Online
    • US flagUnited States
    Respan AI screenshot 1
    Respan AI screenshot 1
    Respan AI screenshot 2
    +1
    Respan AI screenshot 3
    3 alternatives
  9. Tracely icon
     Like

    Open-source LLM observability and evaluation: grade every AI agent trace, cluster the failures, and replay them as regression tests that block the pull request.

    Cost / License

    • Freemium
    • Open Source (MIT)

    Platforms

    • Online
    • Software as a Service (SaaS)
    • Self-Hosted
    • Docker
    • US flagUnited States
    Tracely screenshot 1
    Tracely screenshot 1
    Tracely screenshot 2
    +4
    Tracely screenshot 3
    5 alternatives
  10. Netra icon
     Like

    Netra is the reliability platform for AI agents to observe, evaluate, simulate, and continuously improve every decision your agents make, so you can ship with confidence and catch regressions before your users do.

    Cost / License

    • Freemium
    • Proprietary

    Platforms

    • Online
    • Software as a Service (SaaS)
    • US flagUnited States
    Netra screenshot 1
    Netra screenshot 1
    Netra screenshot 2
    +2
    Netra screenshot 3
    5 alternatives
  11. Agenta icon
     Like

    The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.

    Cost / License

    • Freemium
    • Open Source

    Platforms

    • Online
    • Self-Hosted
    • DE flagGermany
    • European Union flagEU
    Agenta screenshot 1
    Agenta screenshot 1
    Agenta screenshot 2
    +3
    Agenta screenshot 3
    6 alternatives