Запись архива

Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs

arXiv:2608.11232v1 Announce Type: new Abstract: Evaluating LLM coding agents in algorithmic trading is difficult because static benchmarks risk data contamination and numerical backtest outputs require ground truth from actual code execution. We present Backtrader-Bench, a framework with two complementary pipelines. A

Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs
Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs
Journalists Protest against rising violence during march in Mexi | by Knight Foundation | openverse | by-sa

Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs

Что произошло

Источник arXiv cs.CL зафиксировал сигнал: arXiv:2608.11232v1 Announce Type: new
Abstract: Evaluating LLM coding agents in algorithmic trading is difficult because static benchmarks risk data contamination and numerical backtest outputs require ground truth from actual code execution. We present Backtrader-Bench, a framework with two complementary pipelines. A deterministic multiple-choice question (MCQ) pipeline generates questions from backtest configurations across five trading strategies, 33 templates, and three difficulty tiers, with an independent checker that re-derives every answer. A generator-solver filtering pipeline autonomously mines harder questions: a generator writes questions verified by executable code, converts them to MCQs, and discards any that a no-tool solver can answer without code execution. We evaluate 11 models without tools (10 runs each) and four with-tools configurations on a 30-question curated set. Tool-augmented agents reach 90.0% accuracy in a single pass (GPT-5.5 and Opus 4.7), outperforming the best no-tools baselines (73.0%, averaged over 10 runs) by 17 percentage points. On 38 separately mined questions, no-tools accuracy drops further, with half the models falling to roughly random-chance level (25%). Beyond evaluation, the scalable MCQ infrastructure is designed to produce a training corpus for reinforcement learning, with the ultimate goal of building a specialized agent for quantitative trading workflows.

Почему это обсуждают

Для аудитории COMRAD404 это повод проверить, касается ли тема моделей, агентов, промптов, инструментов или разработки с ИИ. Социальный источник сам по себе не является доказательством, поэтому выводы нужно держать осторожными.

Что подтверждено

Punkt Detail
Платформа arxiv
Источник arXiv cs.CL
Проверка https://arxiv.org/

Что проверить дальше

Нужно открыть первичный источник, документацию продукта, GitHub, блог лаборатории или публикацию автора и отделить факт релиза от реакции сообщества.

Источник: arXiv cs.CL – https://arxiv.org/abs/2608.11232; проверка: https://arxiv.org/