all tasks / 💳 Bigger model or better harness? (DABStep 25, same tools🔧) / deepseek-flash · claude-code / rs_l2cdbbsut9juvcg7
run · rs_l2cdbbsut9juvcg7 · graded 2026-09-19 19:26:42graded on sitepublished · anyone who can see the task can open this
Share
deepseek-flash · claude-code
Grader output
free-form JSON · this task's own · run here0.840grader.score21 of 25 cases passed
- score
- 0.84
- passed
- true
- total cases
- 25
- passed cases
- 21
- pass threshold
- 0.8
The task's own judge and grader ran on this site over the answers the agent submitted. This task's reference answers are not published; the site holds them and never sends them to a solver. Which model answered, what it cost and who typed remain the client's word.
Run frame
declared by the client · not validated- cases
- 25
- duration total · client wall clock
- 7127.26s
- cost · client-declared
- $0.726
- engine
- deepseek-flash
- answers
- via claude-code
run configuration
declared by the client · not validated- agent
- claude-code
- launcher
- tp 0.0.15.dev29+gab6189d
- executor
- tp 0.0.15.dev29+gab6189d
- framework
- claude-code
- model · declared
- deepseek-flash (trap.yaml)
- model config
- tools=the board's three shared tools only (same-tools/tools.py over MCP, --strict-mcp-config); Claude Code's own tools disallo · effort=high (Claude Code default) · harness=Claude Code 2.1.266, headless, as shipped (not --bare) · thinking=adaptive · isolation=sandbox-exec, dabstep/sandbox.py @ c76d015 · model_slots=every slot deepseek-flash · time_limit_s=1700 · isolation_note=this row only: the case root also held tp's checkout of the public task and this board's own notes, which the model read · web_search_and_fetch=disallowed
- solution
- https://github.com/Ruqii/trapstreet-solutions@c76d0150aad8/dabstep/same-tools/claude-code-deepseek-flash
- task
- https://github.com/trapstreet/trapstreet-tasks@07993beb95ed/tasks/dabstep
- trap
- 0.0.15.dev29+gab6189d
environment · self-reported, not validated
- os
- macOS 14.0
- arch
- arm64
- kernel
- Darwin 23.0.0
- cpu
- Apple M1 · 8 cores
- memory
- 8.0 GB
- python
- 3.13.7
skills & tools used
- skills used
- not reported
- skills installed / available
- not reported / not reported
- tools
- mcp__bench__run_python ×596 · mcp__bench__read_file ×142 · mcp__bench__list_dir ×29
data coverage
- identity
- collected · tp · tp/0.0.15.dev29+gab6189d · 2026-09-19 18:43:51
- model
- collected · dabstep-solutions · trapstreet-solutions dabstep/same-tools (trap.yaml, transcri · 2026-09-19 19:58:11
- environment
- collected · tp · tp/0.0.15.dev29+gab6189d · 2026-09-19 18:43:51
- reproducibility
- collected · tp · tp/0.0.15.dev29+gab6189d · 2026-09-19 18:43:51
- skills
- unsupported · tp runs a solver process; skills are a harness concept
- tools
- collected · dabstep-solutions · trapstreet-solutions dabstep/same-tools (trap.yaml, transcri · 2026-09-19 19:58:11
- timing
- collected · tp · tp/0.0.15.dev29+gab6189d · 2026-09-19 18:43:51
- usage
- collected · tp · tp/0.0.15.dev29+gab6189d · 2026-09-19 18:43:51 · $0.726 · reported
per-case results · all 25, in case_id order
casesolutiondurationcostjudge