Recent activity audit

arena-hard-auto

Automation harness for Arena-Hard model evaluation and judging.
Last refreshed 2026-08-31 10:36 UTC
← Back to main audit
436d · stale · git
Root
/home/gb001/research_bench/arena-hard-auto
Website URL
Scale
45 files · 14,885 LOC · Small
Status
git on main
Output
Benchmark/evaluation harness with browser QA helpers and data.

Project profile

Purpose: Automation harness for Arena-Hard model evaluation and judging.
Top-level structure: README.md; requirements.txt; BenchBuilder/; config/; data/; leaderboard/; misc/; utils/; +6 more
Top languages: Other (17f/9,953l) Python (15f/3,859l) Markdown (3f/558l) YAML (6f/429l)
Category: Research & benchmarks

Audit trail

Workspace notes