agent_id. The score is a portable 0–1000 integer stored centrally, which means it travels with the agent across every platform and integration mode — whether the agent is gated via the SDK, the proxy, or the REST API. A single score reflects all of an agent’s evaluations regardless of where they originated.
Scoring formula
The score is a weighted sum of four components, each normalised to its maximum contribution:
The maximum possible score is 1000. A brand-new agent with perfect evaluations climbs toward 1000 as it accumulates volume.
Lifecycle states
The score goes through four lifecycle states as evaluations accumulate:The score only becomes visible to external consumers at 50 evaluations. Below that threshold the data is insufficient for a statistically reliable signal, so the dashboard displays the calibration progress instead of a raw number.
Rolling window
Reputation is computed over the last 500 evaluations. When a new evaluation arrives and the window is full, the oldest entry is evicted. This means an agent can recover from a bad period: sustained good behaviour will eventually push earlier failures out of the window.Portability
The score is stored centrally under a stableagent_id per organization. Any integration mode that uses the same agent_id writes to the same window:
- An SDK evaluation in your Node.js service
- A proxy evaluation from Claude Desktop on a developer’s laptop
- A REST API evaluation from your Python test suite
agent_id string, so you can segment agents as finely as you need (e.g. dsp-bidder-prod vs dsp-bidder-staging).
Reading the score
The REST API returns the current reputation inline on every/v1/evaluate response: