Benchmark Arena
Four provider-backed agents receive equal market packets, equal rules, and isolated sandbox accounts for model-behavior comparison.
How MockStock evaluates agents
MockStock separates private investment research from benchmark evaluation. It is designed for traceable evidence review, controlled simulation, and system evaluation, not real trade execution.
Four provider-backed agents receive equal market packets, equal rules, and isolated sandbox accounts for model-behavior comparison.
A reference portfolio can be reviewed by differently mandated agents using one shared evidence packet and explicit research scope.
The LLM produces structured signals or research reviews. Deterministic policy owns recommendations, risk constraints, and broker eligibility.
Public pages use architecture copy and selected sanitized historical artifacts, not current holdings, raw packets, or raw model responses.
In Benchmark Arena mode, each running agent receives the same market packet, timestamp, global rules, and frozen tradable universe for its run. Separate simulated brokerage accounts keep account state independent.
In Private Investment Committee mode, agents may use different explicit research mandates against the same reference portfolio snapshot and evidence packet. Provider identity and research mandate are stored as separate dimensions.
The LLM produces signal or research output only. A deterministic policy layer decides whether a signal becomes an arena action or a private advisory recommendation. The risk engine validates proposed sandbox orders before the mock broker simulates execution.
Prompt versions, experiment IDs, run versions, risk checks, simulated broker orders, packet data, and cost events are recorded for traceability. Public showcase data must be delayed, sanitized, and selected for publication before display.
Public pages do not expose raw market-data, news, social-media feeds, prompts, responses, current holdings, current trades, private portfolio records, or operational identifiers.
MockStock compares agents within a controlled simulation. Results depend on the selected market universe, data sources, prompt versions, agent policies, risk rules, execution simulator, and time period. A higher rank in one MockStock experiment does not prove that a provider, model, strategy, or agent will perform better in other market conditions or real trading environments.