Index / TokenTrim/jev-agent-failure-benchmark
jev-agent-failure-benchmark
Benchmarking Jev (Typesafe.ai) against a strong LLM on the Who&When Pro agent-failure-attribution benchmark (text subset).
Research and EvalsExtractionPythonListed
Repository
Stars0
LanguagePython
LicenseApache-2.0
Created2026-09-17
Last pushed2026-09-17
Officialno
Judgment
Genuine Jev project0.95
Uses Jev at runtime0.90
Is a meta list0.04
Is a reimplementation0.05
Quality scales
Substance2.00 / 3
Docs2.60 / 3
Novelty2.74 / 3
Composite score0.79
Category distribution
Research and Evals1.00
Integrations0.00
Learning0.00
SDKs and Clients0.00
Other0.00
Agent and Dev Tooling0.00
Applications0.00
Games and Simulation0.00
Assigned categoryResearch and Evals
Category probability1.00
Category confidence1.00
PatternExtraction
Pattern confidence0.87
Curation
StatusListed
Reasongate passed
README pickno
Overridenone
Sources
- search:jev typesafe
- search:"typesafe.ai"
- list:hellogumbo/awesome-jev
- list:AnotiaWang/awesome-jev
- search:jev in:name,description,readme created:2026-09-17
Provenance
Modeljev-1.13.0
Question setv2
Judged at2026-09-17 16:34 UTC