Recent activity audit

WildBench

Benchmarking LLMs with challenging tasks from real users in the wild.
Last refreshed 2026-08-31 10:36 UTC
← Back to main audit
760d · stale · git
Root
/home/gb001/research_bench/WildBench
Website URL
Scale
504 files · 8,038,739 LOC · Large
Status
git on main
Output
Benchmark/evaluation harness with leaderboard + eval result corpus.

Project profile

Purpose: Benchmarking LLMs with challenging tasks from real users in the wild.
Top-level structure: README.md; requirements.txt; docs/; eval_results/; evaluation/; leaderboard/; scripts/; src/; +3 more
Top languages: JSON (347f/7,957,147l) CSS (20f/38,367l) JavaScript (28f/33,104l) Python (23f/5,265l)
Category: Research & benchmarks

Audit trail

Workspace notes