reconbench

Indexemblemo-portfoliorun

cc/sonnet

20260818T012743Z-f4e44a

runningdev / untrustedhigh effort · prompt=baseline · matrix · — · — · 2026-08-18T01:27:44.811834+00:00

dev / untrusted. Candidate code ran on the host through the development-only native backend, with the evaluator inside the same trust boundary. Useful evidence, not a publishable benchmark score.How to read this →

01

Verdict

No grade was produced · running

A run with status running is not converted into a score. The manifest and digests below are preserved exactly as recorded so the failure stays auditable.

02

Apparatus

Agent, variant, timing, usage and content digests, as recorded.

Agent

Selector
cc/sonnet
Runtime
claude_code · 2.1.233 (Claude Code)
Model requested
sonnet
Model resolved
Effort
high
Tool profile
baseline

Variant & network

Variant
prompt=baseline · matrix
Network policy
deny
Enforced
false

The declared network policy was not enforced by a sandbox in this run.

Timing & usage

Started
2026-08-18 01:27:44 UTC
Finished
Elapsed
Input tokens
Output tokens
Cost

Content digests

evaluator_environment
55f70d30b5f1a5601563cf20f63ec372ace54a40b2893fd72300a76ff04713b8
grader
oracle
runtime_config
b8f42d18225a47e0c28bb0cb095f136b31668defc61e9d59325e2f3bc1fedce4
submission
task
0733100788cf2f2f084fff23ed8d9862cf151e153bca06e6addd16f4c3302c9e

Digests pin the exact task package, submission tree, oracle, grader and evaluator environment used. They are how a result is re-verified without trusting this page.