all tasks / 🪲 Which code review skill works the best? / baseline-no-skill-claude-opus-4-8 / run_g5lsfvcvsrre1dco
run · run_g5lsfvcvsrre1dco · filed 2026-07-30 15:19:38
Share
baseline-no-skill-claude-opus-4-8
Grader output
free-form JSON · this task's own0.600grader.score6 of 10 cases passed
- score
- 0.600
- passed
- ✓
- total cases
- 10
- passed cases
- 6
- skipped cases
- 0
- pass threshold
- 0.500
- cost total
- $0.081
- latency p95
- 6.94 s
- latency total
- 49.00 s
- latency median
- 4.27 s
score by category
null_deref
0%
off_by_one
100%
broad_except
100%
resource_leak
0%
sql_injection
100%
stale_closure_state
100%
security_missing_auth
0%
cache_invalidate_keyerror
100%
floating_point_truncation
100%
elif_short_circuit_wrong_deny
0%
extras
n_scored=10n_skipped_no_gold=0Fields here are the task author's — trap neither interprets nor validates them.
Run frame
written by trap · always present- cases
- 10
- duration total
- 49.00s
- cost
- $0.081
- engine
- claude-opus-4-8
environment · self-reported, not validated
- os
- macOS 14.0
- arch
- arm64
- kernel
- Darwin 23.0.0
- cpu
- Apple M1 · 8 cores
- memory
- 8.0 GB
- trap
- 0.0.11
per-case results · all 10, in case_id order
casesolutiondurationcostjudge