reconbench

Indexemblemo-portfoliorun

codex/gpt-5.6-sol

20260818T011756Z-6d2875

completeddev / untrustedhigh effort · prompt=baseline · pilot · 8m 27s · — · 2026-08-18T01:17:56.765656+00:00

dev / untrusted. Candidate code ran on the host through the development-only native backend, with the evaluator inside the same trust boundary. Useful evidence, not a publishable benchmark score.How to read this →

01

Verdict

Weights come from the case scoring block, not from this page.

Overall85.7
Score arithmetic
ComponentValueWeightContribution
Build gateinstall + build; a failed gate scores zero overall✓ passedgate
Normalized visualraw 86.6 · floor 0.160.8402× 0.75+0.6301
Functional checksweighted deterministic checks0.9091× 0.25+0.2273
Overall= 0.8574
02

Visual evidence

4 checkpoints · reference against candidate · drag the wipe or fade the onion skin

Fig. 01homebase target · 1440×900

visual 79.5
reference capture for home
reference
candidate capture for home
candidate

pixel 98.0thumbnail 98.6fg-IoU 62.0edge 73.9dimension 100.0

Fig. 02playgroundbase target · 1440×1070

visual 94.4
reference capture for playground
reference
candidate capture for playground
candidate

pixel 99.6thumbnail 99.8fg-IoU 91.8edge 89.6dimension 100.0

Fig. 03home-menu-openstate on home

visual 92.9
reference capture for home-menu-open
reference
candidate capture for home-menu-open
candidate

pixel 98.2thumbnail 98.8fg-IoU 99.4edge 75.1dimension 100.0

Fig. 04playground-closestate on playground

visual 79.5
reference capture for playground-close
reference
candidate capture for playground-close
candidate

pixel 98.0thumbnail 98.6fg-IoU 62.0edge 73.9dimension 100.0

03

Functional checks

Deterministic browser assertions. Selectors stay sealed.

4 / 5 passed · weight 5.0 / 5.5

ResultCheckKindWeightDetail
required-file:submission.json0.5submission.json is missing required keys: routes, status
home-headingtext1.0
menu-expandsattribute_after_click1.5
playground-breadcrumbtext1.0
playground-close-returns-homeclick_text_change1.5
04

Apparatus

Agent, variant, timing, usage and content digests, as recorded.

Agent

Selector
codex/gpt-5.6-sol
Runtime
codex · codex-cli 0.147.0
Model requested
gpt-5.6-sol
Model resolved
Effort
high
Tool profile
baseline

Variant & network

Variant
prompt=baseline · pilot
Network policy
deny
Enforced
true

Timing & usage

Started
2026-08-18 01:17:56 UTC
Finished
2026-08-18 01:26:23 UTC
Elapsed
8m 27s
Input tokens
1.65M
Output tokens
21.9k
Cost

Content digests

evaluator_environment
55f70d30b5f1a5601563cf20f63ec372ace54a40b2893fd72300a76ff04713b8
grader
9d083a514d47783415ea76b5032d09e0988855c7801293225c9cb03883a1b0f2
oracle
06ccacf5f2c9c73db684278f2c0fcb34a4b9b84c606989d09a62ae510ea1a340
runtime_config
8cdefe81278603d0a8dfcc1030156c79de4110d782608d3795630ab86c28982b
submission
d2b4eaabca6236cfeb99e3683575d92644bf907711fe49fb90e0ec5c7da17a06
task
0733100788cf2f2f084fff23ed8d9862cf151e153bca06e6addd16f4c3302c9e

Digests pin the exact task package, submission tree, oracle, grader and evaluator environment used. They are how a result is re-verified without trusting this page.