Add IntelligenceLab/Long-Horizon-Terminal-Bench to the Benchmark allow-list

Hi HF team,

I’d like to register IntelligenceLab/Long-Horizon-Terminal-Bench as a Benchmark so its leaderboard can aggregate community evaluation results across the Hub.

Dataset: IntelligenceLab/Long-Horizon-Terminal-Bench · Datasets at Hugging Face
Paper: [2607.08964] Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

LHTB is a 46-task benchmark testing how well LLM agents sustain useful work in a containerized terminal over hundreds of steps, using hidden, rebuild-from-artifact verifiers rather than self-reported progress. Tasks run via the Harbor / Terminal-Bench 2.0 harness.

eval.yaml is already in the repo root (evaluation_framework: harbor) and validated on push.

Could you add this to the Benchmark allow-list? Thanks!

I’m not HF Staff or so, but, for now, see also this around benchmark registration process: Add haifan-gong/CTGroundBench to the Benchmark allow-list - #2 by John6666