Can your model identify animals?
ranked by score ↓Source
Paste as source: in your trap.yaml
git+https://github.com/alfred-ruqilabs/identify-the-animal@efaa614a016be404b18ea64f45ee343f8644a859Share
identify-the-animal
This repository contains a vision test where the job is simple: look at a wildlife photo and say which animal species it shows.
20 cases
Each case feeds files from inputs/<id>/ to the solution, expects files in expected/<id>/, and is scored by judge.py then aggregated by grader.py.
cases (20)
▸zebra_01Identify the species in this camera trap photo
input
document.jpg· 267.3 KBview on GitHub

question.txt
This image is from a wildlife camera trap in Serengeti National Park. What is the main animal species visible in the photograph?
Choose from: buffalo, cheetah, elephant, giraffe, hyena, leopard, lion, warthog, wildebeest, zebra
Answer with exactly one word — the species name from the list. Do not hedge or explain.
expected output
answer.json
{
"id": "zebra_01",
"answer": "zebra",
"type": "species_classification",
"matchers": [
{
"kind": "leading_word",
"value": "zebra"
},
{
"kind": "no_hedge"
}
],
"category": "zebra",
"difficulty": "easy",
"source_dataset_label": "zebra",
"source_capture_event": "ASG000f1a4",
"source_url_path": "S4/O07/O07_R1/S4_O07_R1_IMAG0514.JPG"
}Scored by judge.py — see Scoring logic below for the full rule.
▸zebra_02Identify the species in this camera trap photo
input
document.jpg· 144.7 KBview on GitHub

question.txt
This image is from a wildlife camera trap in Serengeti National Park. What is the main animal species visible in the photograph?
Choose from: buffalo, cheetah, elephant, giraffe, hyena, leopard, lion, warthog, wildebeest, zebra
Answer with exactly one word — the species name from the list. Do not hedge or explain.
expected output
answer.json
{
"id": "zebra_02",
"answer": "zebra",
"type": "species_classification",
"matchers": [
{
"kind": "leading_word",
"value": "zebra"
},
{
"kind": "no_hedge"
}
],
"category": "zebra",
"difficulty": "easy",
"source_dataset_label": "zebra",
"source_capture_event": "ASG000blpa",
"source_url_path": "S4/Q05/Q05_R1/S4_Q05_R1_IMAG0661.JPG"
}Scored by judge.py — see Scoring logic below for the full rule.
▸wildebeest_01Identify the species in this camera trap photo
input
document.jpg· 363.0 KBview on GitHub

question.txt
This image is from a wildlife camera trap in Serengeti National Park. What is the main animal species visible in the photograph?
Choose from: buffalo, cheetah, elephant, giraffe, hyena, leopard, lion, warthog, wildebeest, zebra
Answer with exactly one word — the species name from the list. Do not hedge or explain.
expected output
answer.json
{
"id": "wildebeest_01",
"answer": "wildebeest",
"type": "species_classification",
"matchers": [
{
"kind": "leading_word",
"value": "wildebeest"
},
{
"kind": "no_hedge"
}
],
"category": "wildebeest",
"difficulty": "easy",
"source_dataset_label": "wildebeest",
"source_capture_event": "ASG000ckq2",
"source_url_path": "S4/U09/U09_R2/S4_U09_R2_IMAG0691.JPG"
}Scored by judge.py — see Scoring logic below for the full rule.
▸wildebeest_02Identify the species in this camera trap photo
input
document.jpg· 272.6 KBview on GitHub

question.txt
This image is from a wildlife camera trap in Serengeti National Park. What is the main animal species visible in the photograph?
Choose from: buffalo, cheetah, elephant, giraffe, hyena, leopard, lion, warthog, wildebeest, zebra
Answer with exactly one word — the species name from the list. Do not hedge or explain.
expected output
answer.json
{
"id": "wildebeest_02",
"answer": "wildebeest",
"type": "species_classification",
"matchers": [
{
"kind": "leading_word",
"value": "wildebeest"
},
{
"kind": "no_hedge"
}
],
"category": "wildebeest",
"difficulty": "easy",
"source_dataset_label": "wildebeest",
"source_capture_event": "ASG000cju3",
"source_url_path": "S4/C13/C13_R1/S4_C13_R1_IMAG1078.JPG"
}Scored by judge.py — see Scoring logic below for the full rule.
▸buffalo_01Identify the species in this camera trap photo
input
document.jpg· 146.7 KBview on GitHub

question.txt
This image is from a wildlife camera trap in Serengeti National Park. What is the main animal species visible in the photograph?
Choose from: buffalo, cheetah, elephant, giraffe, hyena, leopard, lion, warthog, wildebeest, zebra
Answer with exactly one word — the species name from the list. Do not hedge or explain.
expected output
answer.json
{
"id": "buffalo_01",
"answer": "buffalo",
"type": "species_classification",
"matchers": [
{
"kind": "leading_word",
"value": "buffalo"
},
{
"kind": "no_hedge"
}
],
"category": "buffalo",
"difficulty": "easy",
"source_dataset_label": "buffalo",
"source_capture_event": "ASG000biq2",
"source_url_path": "S4/L06/L06_R2/S4_L06_R2_IMAG0096.JPG"
}Scored by judge.py — see Scoring logic below for the full rule.
▸buffalo_02Identify the species in this camera trap photo
input
document.jpg· 168.6 KBview on GitHub

question.txt
This image is from a wildlife camera trap in Serengeti National Park. What is the main animal species visible in the photograph?
Choose from: buffalo, cheetah, elephant, giraffe, hyena, leopard, lion, warthog, wildebeest, zebra
Answer with exactly one word — the species name from the list. Do not hedge or explain.
expected output
answer.json
{
"id": "buffalo_02",
"answer": "buffalo",
"type": "species_classification",
"matchers": [
{
"kind": "leading_word",
"value": "buffalo"
},
{
"kind": "no_hedge"
}
],
"category": "buffalo",
"difficulty": "easy",
"source_dataset_label": "buffalo",
"source_capture_event": "ASG000ekfb",
"source_url_path": "S4/G04/G04_R1/S4_G04_R1_IMAG0860.JPG"
}Scored by judge.py — see Scoring logic below for the full rule.
▸elephant_01Identify the species in this camera trap photo
input
document.jpg· 77.4 KBview on GitHub

question.txt
This image is from a wildlife camera trap in Serengeti National Park. What is the main animal species visible in the photograph?
Choose from: buffalo, cheetah, elephant, giraffe, hyena, leopard, lion, warthog, wildebeest, zebra
Answer with exactly one word — the species name from the list. Do not hedge or explain.
expected output
answer.json
{
"id": "elephant_01",
"answer": "elephant",
"type": "species_classification",
"matchers": [
{
"kind": "leading_word",
"value": "elephant"
},
{
"kind": "no_hedge"
}
],
"category": "elephant",
"difficulty": "easy",
"source_dataset_label": "elephant",
"source_capture_event": "ASG000bjyt",
"source_url_path": "S4/Q08/Q08_R2/S4_Q08_R2_IMAG1052.JPG"
}Scored by judge.py — see Scoring logic below for the full rule.
▸elephant_02Identify the species in this camera trap photo
input
document.jpg· 293.8 KBview on GitHub

question.txt
This image is from a wildlife camera trap in Serengeti National Park. What is the main animal species visible in the photograph?
Choose from: buffalo, cheetah, elephant, giraffe, hyena, leopard, lion, warthog, wildebeest, zebra
Answer with exactly one word — the species name from the list. Do not hedge or explain.
expected output
answer.json
{
"id": "elephant_02",
"answer": "elephant",
"type": "species_classification",
"matchers": [
{
"kind": "leading_word",
"value": "elephant"
},
{
"kind": "no_hedge"
}
],
"category": "elephant",
"difficulty": "easy",
"source_dataset_label": "elephant",
"source_capture_event": "ASG000f8sd",
"source_url_path": "S4/B06/B06_R2/S4_B06_R2_IMAG1297.JPG"
}Scored by judge.py — see Scoring logic below for the full rule.
▸giraffe_01Identify the species in this camera trap photo
input
document.jpg· 209.9 KBview on GitHub

question.txt
This image is from a wildlife camera trap in Serengeti National Park. What is the main animal species visible in the photograph?
Choose from: buffalo, cheetah, elephant, giraffe, hyena, leopard, lion, warthog, wildebeest, zebra
Answer with exactly one word — the species name from the list. Do not hedge or explain.
expected output
answer.json
{
"id": "giraffe_01",
"answer": "giraffe",
"type": "species_classification",
"matchers": [
{
"kind": "leading_word",
"value": "giraffe"
},
{
"kind": "no_hedge"
}
],
"category": "giraffe",
"difficulty": "easy",
"source_dataset_label": "giraffe",
"source_capture_event": "ASG000bbcm",
"source_url_path": "S4/E05/E05_R2/S4_E05_R2_IMAG1259.JPG"
}Scored by judge.py — see Scoring logic below for the full rule.
▸giraffe_02Identify the species in this camera trap photo
input
document.jpg· 218.9 KBview on GitHub

question.txt
This image is from a wildlife camera trap in Serengeti National Park. What is the main animal species visible in the photograph?
Choose from: buffalo, cheetah, elephant, giraffe, hyena, leopard, lion, warthog, wildebeest, zebra
Answer with exactly one word — the species name from the list. Do not hedge or explain.
expected output
answer.json
{
"id": "giraffe_02",
"answer": "giraffe",
"type": "species_classification",
"matchers": [
{
"kind": "leading_word",
"value": "giraffe"
},
{
"kind": "no_hedge"
}
],
"category": "giraffe",
"difficulty": "easy",
"source_dataset_label": "giraffe",
"source_capture_event": "ASG000bf1k",
"source_url_path": "S4/K07/K07_R2/S4_K07_R2_IMAG2488.JPG"
}Scored by judge.py — see Scoring logic below for the full rule.
▸warthog_01Identify the species in this camera trap photo
input
document.jpg· 469.3 KBview on GitHub

question.txt
This image is from a wildlife camera trap in Serengeti National Park. What is the main animal species visible in the photograph?
Choose from: buffalo, cheetah, elephant, giraffe, hyena, leopard, lion, warthog, wildebeest, zebra
Answer with exactly one word — the species name from the list. Do not hedge or explain.
expected output
answer.json
{
"id": "warthog_01",
"answer": "warthog",
"type": "species_classification",
"matchers": [
{
"kind": "leading_word",
"value": "warthog"
},
{
"kind": "no_hedge"
}
],
"category": "warthog",
"difficulty": "medium",
"source_dataset_label": "warthog",
"source_capture_event": "ASG000dim1",
"source_url_path": "S4/D04/D04_R2/S4_D04_R2_IMAG0211.JPG"
}Scored by judge.py — see Scoring logic below for the full rule.
▸warthog_02Identify the species in this camera trap photo
input
document.jpg· 231.5 KBview on GitHub

question.txt
This image is from a wildlife camera trap in Serengeti National Park. What is the main animal species visible in the photograph?
Choose from: buffalo, cheetah, elephant, giraffe, hyena, leopard, lion, warthog, wildebeest, zebra
Answer with exactly one word — the species name from the list. Do not hedge or explain.
expected output
answer.json
{
"id": "warthog_02",
"answer": "warthog",
"type": "species_classification",
"matchers": [
{
"kind": "leading_word",
"value": "warthog"
},
{
"kind": "no_hedge"
}
],
"category": "warthog",
"difficulty": "medium",
"source_dataset_label": "warthog",
"source_capture_event": "ASG000dywy",
"source_url_path": "S4/I02/I02_R2/S4_I02_R2_IMAG0291.JPG"
}Scored by judge.py — see Scoring logic below for the full rule.
▸lion_01Identify the species in this camera trap photo
input
document.jpg· 156.6 KBview on GitHub

question.txt
This image is from a wildlife camera trap in Serengeti National Park. What is the main animal species visible in the photograph?
Choose from: buffalo, cheetah, elephant, giraffe, hyena, leopard, lion, warthog, wildebeest, zebra
Answer with exactly one word — the species name from the list. Do not hedge or explain.
expected output
answer.json
{
"id": "lion_01",
"answer": "lion",
"type": "species_classification",
"matchers": [
{
"kind": "leading_word",
"value": "lion"
},
{
"kind": "no_hedge"
}
],
"category": "lion",
"difficulty": "medium",
"source_dataset_label": "lionfemale",
"source_capture_event": "ASG000c9do",
"source_url_path": "S4/I13/I13_R1/S4_I13_R1_IMAG0373.JPG"
}Scored by judge.py — see Scoring logic below for the full rule.
▸lion_02Identify the species in this camera trap photo
input
document.jpg· 302.4 KBview on GitHub

question.txt
This image is from a wildlife camera trap in Serengeti National Park. What is the main animal species visible in the photograph?
Choose from: buffalo, cheetah, elephant, giraffe, hyena, leopard, lion, warthog, wildebeest, zebra
Answer with exactly one word — the species name from the list. Do not hedge or explain.
expected output
answer.json
{
"id": "lion_02",
"answer": "lion",
"type": "species_classification",
"matchers": [
{
"kind": "leading_word",
"value": "lion"
},
{
"kind": "no_hedge"
}
],
"category": "lion",
"difficulty": "medium",
"source_dataset_label": "lionfemale",
"source_capture_event": "ASG000e8zf",
"source_url_path": "S4/J12/J12_R1/S4_J12_R1_IMAG0297.JPG"
}Scored by judge.py — see Scoring logic below for the full rule.
▸leopard_01Identify the species in this camera trap photo
input
document.jpg· 278.6 KBview on GitHub

question.txt
This image is from a wildlife camera trap in Serengeti National Park. What is the main animal species visible in the photograph?
Choose from: buffalo, cheetah, elephant, giraffe, hyena, leopard, lion, warthog, wildebeest, zebra
Answer with exactly one word — the species name from the list. Do not hedge or explain.
expected output
answer.json
{
"id": "leopard_01",
"answer": "leopard",
"type": "species_classification",
"matchers": [
{
"kind": "leading_word",
"value": "leopard"
},
{
"kind": "no_hedge"
}
],
"category": "leopard",
"difficulty": "hard",
"source_dataset_label": "leopard",
"source_capture_event": "ASG000fasd",
"source_url_path": "S4/U11/U11_R2/S4_U11_R2_IMAG4123.JPG"
}Scored by judge.py — see Scoring logic below for the full rule.
▸leopard_02Identify the species in this camera trap photo
input
document.jpg· 219.5 KBview on GitHub

question.txt
This image is from a wildlife camera trap in Serengeti National Park. What is the main animal species visible in the photograph?
Choose from: buffalo, cheetah, elephant, giraffe, hyena, leopard, lion, warthog, wildebeest, zebra
Answer with exactly one word — the species name from the list. Do not hedge or explain.
expected output
answer.json
{
"id": "leopard_02",
"answer": "leopard",
"type": "species_classification",
"matchers": [
{
"kind": "leading_word",
"value": "leopard"
},
{
"kind": "no_hedge"
}
],
"category": "leopard",
"difficulty": "hard",
"source_dataset_label": "leopard",
"source_capture_event": "ASG000e8r6",
"source_url_path": "S4/B05/B05_R1/S4_B05_R1_IMAG0151.JPG"
}Scored by judge.py — see Scoring logic below for the full rule.
▸cheetah_01Identify the species in this camera trap photo
input
document.jpg· 201.8 KBview on GitHub

question.txt
This image is from a wildlife camera trap in Serengeti National Park. What is the main animal species visible in the photograph?
Choose from: buffalo, cheetah, elephant, giraffe, hyena, leopard, lion, warthog, wildebeest, zebra
Answer with exactly one word — the species name from the list. Do not hedge or explain.
expected output
answer.json
{
"id": "cheetah_01",
"answer": "cheetah",
"type": "species_classification",
"matchers": [
{
"kind": "leading_word",
"value": "cheetah"
},
{
"kind": "no_hedge"
}
],
"category": "cheetah",
"difficulty": "hard",
"source_dataset_label": "cheetah",
"source_capture_event": "ASG000cxls",
"source_url_path": "S4/L08/L08_R2/S4_L08_R2_IMAG2410.JPG"
}Scored by judge.py — see Scoring logic below for the full rule.
▸cheetah_02Identify the species in this camera trap photo
input
document.jpg· 199.6 KBview on GitHub

question.txt
This image is from a wildlife camera trap in Serengeti National Park. What is the main animal species visible in the photograph?
Choose from: buffalo, cheetah, elephant, giraffe, hyena, leopard, lion, warthog, wildebeest, zebra
Answer with exactly one word — the species name from the list. Do not hedge or explain.
expected output
answer.json
{
"id": "cheetah_02",
"answer": "cheetah",
"type": "species_classification",
"matchers": [
{
"kind": "leading_word",
"value": "cheetah"
},
{
"kind": "no_hedge"
}
],
"category": "cheetah",
"difficulty": "hard",
"source_dataset_label": "cheetah",
"source_capture_event": "ASG000d5bt",
"source_url_path": "S4/D06/D06_R2/S4_D06_R2_IMAG2845.JPG"
}Scored by judge.py — see Scoring logic below for the full rule.
▸hyena_01Identify the species in this camera trap photo
input
document.jpg· 194.1 KBview on GitHub

question.txt
This image is from a wildlife camera trap in Serengeti National Park. What is the main animal species visible in the photograph?
Choose from: buffalo, cheetah, elephant, giraffe, hyena, leopard, lion, warthog, wildebeest, zebra
Answer with exactly one word — the species name from the list. Do not hedge or explain.
expected output
answer.json
{
"id": "hyena_01",
"answer": "hyena",
"type": "species_classification",
"matchers": [
{
"kind": "leading_word",
"value": "hyena"
},
{
"kind": "no_hedge"
}
],
"category": "hyena",
"difficulty": "medium",
"source_dataset_label": "hyenaspotted",
"source_capture_event": "ASG000bv69",
"source_url_path": "S4/H09/H09_R2/S4_H09_R2_IMAG0382.JPG"
}Scored by judge.py — see Scoring logic below for the full rule.
▸hyena_02Identify the species in this camera trap photo
input
document.jpg· 219.0 KBview on GitHub

question.txt
This image is from a wildlife camera trap in Serengeti National Park. What is the main animal species visible in the photograph?
Choose from: buffalo, cheetah, elephant, giraffe, hyena, leopard, lion, warthog, wildebeest, zebra
Answer with exactly one word — the species name from the list. Do not hedge or explain.
expected output
answer.json
{
"id": "hyena_02",
"answer": "hyena",
"type": "species_classification",
"matchers": [
{
"kind": "leading_word",
"value": "hyena"
},
{
"kind": "no_hedge"
}
],
"category": "hyena",
"difficulty": "medium",
"source_dataset_label": "hyenaspotted",
"source_capture_event": "ASG000c48m",
"source_url_path": "S4/T13/T13_R2/S4_T13_R2_IMAG0267.JPG"
}Scored by judge.py — see Scoring logic below for the full rule.
scoring logic
judge.py runs once per case and prints a score per case. grader.py runs once at the end and folds case scores into a run-level summary. Without grader.py, the run's score is simply the average of case scores.
▸judge.py358 lines · view on GitHub
"""Per-case judge for the tenancy_agreement task — harsh by design.
Reads the agent's stdout (plain text OR JSON `{"answer": "..."}`) and applies
matchers declared in expected/{case_id}/answer.json. A case scores 1.0 only if
ALL matchers pass — partial credit is intentionally not offered. The whole
point of this task is to expose agents that hedge, miss clauses, or skip parts
of multi-part questions; lenient grading would defeat that.
Matcher kinds supported:
- numeric {"kind":"numeric","value":1234.5,"tolerance":0.01}
Passes if ANY number in the answer matches. Use for
show-your-working questions where the model walks
through arithmetic before stating the total.
- leading_numeric {"kind":"leading_numeric","value":1234.5,"tolerance":0.01}
The FIRST number in the answer must match. Use for
simple extraction questions where listing decoy
numbers should not count as a pass.
- regex_required {"kind":"regex_required","pattern":"...","flags":"i"}
Pattern must match (re.search). Default flags = i.
- leading_word {"kind":"leading_word","value":"yes"}
First alphanumeric token must equal value (case-insens),
after stripping common prefixes like "Answer:" or
markdown bold. Forces the model to commit, not hedge.
- keywords_all {"kind":"keywords_all","values":["a","b"]}
Every value must appear (case-insens substring).
- keywords_any {"kind":"keywords_any","values":["a","b"]}
At least one value must appear (case-insens substring).
- keywords_any_word {"kind":"keywords_any_word","values":["ICE","BOE"]}
At least one value must appear as a whole word (\b...\b,
case-insens). Use for short acronyms that would
false-positive as substrings (ICE in "price", BOE in
"Boeing").
- no_hedge {"kind":"no_hedge"}
Reject answers that visibly punt the question, e.g.
"I cannot determine", "unclear from the document",
"I don't have access", "as an AI", etc.
- min_words {"kind":"min_words","value":5}
Reject one-word answers when the question asked for
reasoning/explanation.
Fallback (when no `matchers` provided):
Substring match of `answer` (and any `accepted` variants) against the
normalised agent output. Lenient but kept for cases that haven't been
hardened yet (e.g. scenario_* cases without a curated gold).
Outputs JSON on stdout — trap stores it as CaseResult.metrics. The grader
reads `metrics.score` plus category/difficulty/reason for the report.
"""
from __future__ import annotations
import json
import os
import re
from pathlib import Path
from typing import Any
HEDGE_PHRASES = [
"i cannot", "i can't", "i am unable", "i'm unable",
"i don't have access", "i do not have access",
"as an ai", "as a language model",
"cannot determine", "unable to determine",
"unclear from the document", "not clear from the document",
"i don't know", "i do not know",
"insufficient information", "not enough information",
"i'm not sure", "i am not sure",
]
NUMBER_RE = re.compile(r"-?\d[\d,]*(?:\.\d+)?")
def normalise(s: str) -> str:
return re.sub(r"\s+", " ", s).strip().lower()
def extract_agent_answer(stdout: str) -> str:
"""Accept JSON {"answer": "..."} or plain text. Strip surrounding whitespace."""
stdout = stdout.strip()
if not stdout:
return ""
try:
obj = json.loads(stdout)
if isinstance(obj, dict) and "answer" in obj:
return str(obj["answer"])
except json.JSONDecodeError:
pass
return stdout
def parse_numeric(s: str) -> float | None:
"""Extract the first plausible number from `s`. £, $, commas, spaces stripped."""
nums = parse_all_numerics(s)
return nums[0] if nums else None
def parse_all_numerics(s: str) -> list[float]:
"""Extract ALL plausible numbers from `s`. Used to match agents that show
working (e.g. "£1,950 × 12 + ... = £77,400" — we want to find 77400)."""
if not s:
return []
cleaned = s.replace("£", "").replace("$", "").replace(",", "")
out: list[float] = []
for m in NUMBER_RE.finditer(cleaned):
try:
out.append(float(m.group(0).replace(",", "")))
except ValueError:
continue
return out
_LEADING_LABEL_RE = re.compile(
r"^\s*(?:answer|a|response|reply)[\s*_`]*:\s*", re.IGNORECASE,
)
_LEADING_NOISE_RE = re.compile(r"^[\s*_`#>\-]+")
def leading_word(s: str) -> str:
"""First alpha token, after stripping markdown noise and labels like
"Answer:" / "**Answer**:" / "> ". Lets models prefix their commit with
a natural label without auto-failing the case."""
s = _LEADING_NOISE_RE.sub("", s)
s = _LEADING_LABEL_RE.sub("", s)
s = _LEADING_NOISE_RE.sub("", s)
m = re.search(r"[a-zA-Z]+", s)
return m.group(0).lower() if m else ""
# --- Matcher implementations ----------------------------------------------
def m_numeric(answer: str, spec: dict) -> tuple[bool, str]:
"""Pass if ANY number in the answer matches the target within tolerance.
This lets models that show working ("1950 × 12 + 2100 × 12 = 77400") pass
as long as the right number appears somewhere — exposing the actual answer
is what matters, not whether the model led with it. For simple extraction
where listing decoys should NOT pass, use `leading_numeric` instead."""
nums = parse_all_numerics(answer)
if not nums:
return False, "no number found in answer"
target = float(spec["value"])
tol = float(spec.get("tolerance", 0.01))
for n in nums:
if abs(n - target) <= tol:
return True, f"numeric ok (matched {n} of {nums} against target={target} tol={tol})"
return False, f"numeric mismatch (numbers found={nums} target={target} tol={tol})"
def m_leading_numeric(answer: str, spec: dict) -> tuple[bool, str]:
"""First number in the answer must match within tolerance. Rejects
decoy-number dumps like "rent 1950, deposit 2250, rent yr2 2100"
where the target appears but isn't the committed answer."""
nums = parse_all_numerics(answer)
if not nums:
return False, "no number found in answer"
target = float(spec["value"])
tol = float(spec.get("tolerance", 0.01))
if abs(nums[0] - target) <= tol:
return True, f"leading number ok ({nums[0]} == target {target} tol {tol})"
return False, f"leading number {nums[0]} ≠ target {target} (other numbers in answer: {nums[1:]})"
def m_regex_required(answer: str, spec: dict) -> tuple[bool, str]:
flags = 0
if "i" in spec.get("flags", "i"):
flags |= re.IGNORECASE
if re.search(spec["pattern"], answer, flags):
return True, f"regex matched"
return False, f"regex {spec['pattern']!r} did not match"
def m_leading_word(answer: str, spec: dict) -> tuple[bool, str]:
got = leading_word(answer)
want = str(spec["value"]).lower()
if got == want:
return True, f"leading word ok ({got!r})"
return False, f"leading word {got!r} ≠ required {want!r}"
def m_keywords_all(answer: str, spec: dict) -> tuple[bool, str]:
norm = normalise(answer)
missing = [v for v in spec["values"] if v.lower() not in norm]
if missing:
return False, f"missing required keyword(s): {missing}"
return True, "all keywords present"
def m_keywords_any(answer: str, spec: dict) -> tuple[bool, str]:
norm = normalise(answer)
if any(v.lower() in norm for v in spec["values"]):
return True, "at least one keyword present"
return False, f"none of {spec['values']} present"
def m_keywords_any_word(answer: str, spec: dict) -> tuple[bool, str]:
"""Whole-word variant of keywords_any — wraps each value in \\b...\\b so
short acronyms (ICE, BOE) don't false-match inside "price", "Boeing", etc."""
for v in spec["values"]:
if re.search(rf"\b{re.escape(v)}\b", answer, re.IGNORECASE):
return True, f"whole-word match: {v!r}"
return False, f"none of {spec['values']} matched as whole word"
def m_no_hedge(answer: str, spec: dict) -> tuple[bool, str]:
norm = normalise(answer)
for phrase in HEDGE_PHRASES:
if phrase in norm:
return False, f"hedge phrase detected: {phrase!r}"
return True, "no hedge phrases"
def m_min_words(answer: str, spec: dict) -> tuple[bool, str]:
count = len(re.findall(r"\S+", answer))
want = int(spec["value"])
if count >= want:
return True, f"word count ok ({count} ≥ {want})"
return False, f"too short ({count} < {want})"
MATCHERS = {
"numeric": m_numeric,
"leading_numeric": m_leading_numeric,
"regex_required": m_regex_required,
"leading_word": m_leading_word,
"keywords_all": m_keywords_all,
"keywords_any": m_keywords_any,
"keywords_any_word": m_keywords_any_word,
"no_hedge": m_no_hedge,
"min_words": m_min_words,
}
def run_matchers(answer: str, matchers: list[dict]) -> tuple[float, list[dict]]:
"""Run all matchers; all must pass. Returns (score, per-matcher results)."""
results = []
all_ok = True
for spec in matchers:
kind = spec.get("kind")
fn = MATCHERS.get(kind)
if fn is None:
results.append({"kind": kind, "pass": False, "reason": f"unknown matcher kind: {kind!r}"})
all_ok = False
continue
ok, reason = fn(answer, spec)
results.append({"kind": kind, "pass": ok, "reason": reason})
if not ok:
all_ok = False
return (1.0 if all_ok else 0.0), results
def fallback_substring(answer: str, expected: dict) -> tuple[float, str]:
"""Lenient substring match when no matchers defined. Used for scenarios
that don't have a curated gold yet — they shouldn't fail builds outright,
but they also shouldn't claim a passing score from nothing."""
targets = [t for t in [expected.get("answer"), *(expected.get("accepted") or [])] if t]
if not targets:
return 0.0, "no gold answer set (skip-equivalent)"
norm = normalise(answer)
hit = next((t for t in targets if normalise(t) in norm), None)
if hit:
return 1.0, f"substring match ({hit!r})"
return 0.0, f"no substring match against {targets}"
# --- Main ------------------------------------------------------------------
def main() -> None:
payload = json.loads(os.environ["TRAPTASK_PAYLOAD"])
stdout = Path(payload["outputs"]["case_stdout"]).read_text()
exit_code = json.loads(Path(payload["outputs"]["case_meta.json"]).read_text())["exit_code"]
expected = json.loads(Path(payload["expected"]["answer.json"]).read_text())
# Pick up usage.json if the solution captured it (Sonnet + caching runs)
usage_record: dict[str, Any] = {}
usage_path = payload["outputs"].get("usage.json")
if usage_path and Path(usage_path).exists():
try:
usage_record = json.loads(Path(usage_path).read_text())
except json.JSONDecodeError:
usage_record = {}
agent_answer = extract_agent_answer(stdout)
# Solution crashed → hard fail.
if exit_code != 0:
out: dict[str, Any] = {
"score": 0.0,
"reason": f"solution exited {exit_code}",
"agent_answer": agent_answer,
"id": expected.get("id"),
"category": expected.get("category"),
"difficulty": expected.get("difficulty"),
**usage_record,
}
print(json.dumps(out))
return
# Empty stdout → hard fail (silently passing the test is the worst outcome).
if not agent_answer:
out = {
"score": 0.0,
"reason": "agent produced no answer",
"agent_answer": "",
"id": expected.get("id"),
"category": expected.get("category"),
"difficulty": expected.get("difficulty"),
**usage_record,
}
print(json.dumps(out))
return
matchers = expected.get("matchers")
if matchers:
score, matcher_results = run_matchers(agent_answer, matchers)
out = {
"score": score,
"matcher_results": matcher_results,
"agent_answer": agent_answer,
"expected_answer": expected.get("answer"),
"id": expected.get("id"),
"type": expected.get("type"),
"category": expected.get("category"),
"difficulty": expected.get("difficulty"),
**usage_record,
}
else:
score, reason = fallback_substring(agent_answer, expected)
# If there's no gold and no matchers, surface score=None so the grader
# can flag it as "not yet curated" rather than mark the agent failed.
if expected.get("answer") is None:
out = {
"score": None,
"reason": "no curated gold yet (case not gradeable)",
"agent_answer": agent_answer,
"id": expected.get("id"),
"type": expected.get("type"),
"category": expected.get("category"),
"difficulty": expected.get("difficulty"),
**usage_record,
}
else:
out = {
"score": score,
"reason": reason,
"agent_answer": agent_answer,
"expected_answer": expected.get("answer"),
"id": expected.get("id"),
"type": expected.get("type"),
"category": expected.get("category"),
"difficulty": expected.get("difficulty"),
**usage_record,
}
print(json.dumps(out))
if __name__ == "__main__":
main()
▸grader.py81 lines · view on GitHub
"""Overall grader for the tenancy_agreement task.
Aggregates per-case judge results into a run-level verdict. Emits JSON to stdout —
trap stores it as GraderResult.metrics. Convention: include `passed` (bool) and
`score` (float) so the reporter can render them.
Pass threshold defaults to 80% accuracy; tweak below.
"""
from __future__ import annotations
import json
import os
from collections import Counter
PASS_THRESHOLD = 0.80
def main() -> None:
cases = json.loads(os.environ["TRAPTASK_PAYLOAD"])
scored = [c for c in cases if c.get("metrics") and c["metrics"].get("score") is not None]
skipped = [c for c in cases if not c.get("metrics") or c["metrics"].get("score") is None]
if scored:
accuracy = sum(c["metrics"]["score"] for c in scored) / len(scored)
else:
accuracy = 0.0
# Break out accuracy by category (the judge tags each case with its category).
by_category_score: Counter[str] = Counter()
by_category_total: Counter[str] = Counter()
for c in scored:
cat = c["metrics"].get("category")
if cat:
by_category_total[cat] += 1
by_category_score[cat] += c["metrics"]["score"]
by_category_pct = {
k: round(by_category_score[k] / by_category_total[k], 3)
for k in by_category_total
}
passed = bool(scored) and accuracy >= PASS_THRESHOLD
# Latency stats — trap records `duration` (seconds) per case. The leaderboard
# displays median latency. Round-trip to ms for the JSON contract.
durations = [c.get("duration", 0.0) for c in cases if c.get("duration") is not None]
if durations:
ds = sorted(durations)
latency_ms_median = round(ds[len(ds) // 2] * 1000, 1)
latency_ms_p95 = round(ds[int(0.95 * len(ds))] * 1000, 1) if len(ds) > 1 else latency_ms_median
latency_ms_total = round(sum(ds) * 1000, 1)
else:
latency_ms_median = latency_ms_p95 = latency_ms_total = 0.0
# Cost — sum per-case usd_cost if the solution captured usage; otherwise
# leave it None and let the regrade/submit script stamp a known total.
case_costs = [c["metrics"].get("usd_cost") for c in scored if isinstance(c.get("metrics"), dict)]
cost_usd_total = round(sum(x for x in case_costs if x is not None), 4) if any(x is not None for x in case_costs) else None
n_passed = sum(1 for c in scored if c["metrics"]["score"] == 1.0)
print(json.dumps({
"passed": passed,
"score": round(accuracy, 3),
"n_passed": n_passed,
"n_total": len(cases),
"n_scored": len(scored),
"n_skipped_no_gold": len(skipped),
"threshold": PASS_THRESHOLD,
"by_category": by_category_pct,
"latency_ms_median": latency_ms_median,
"latency_ms_p95": latency_ms_p95,
"latency_ms_total": latency_ms_total,
"cost_usd_total": cost_usd_total,
}))
if __name__ == "__main__":
main()