ROBCO INDUSTRIES (TM) TERMLINK PROTOCOL :: NIGHT CITY INVESTING SUBNETACCESSING DOC ARCHIVE :: GROK-AND-CLAUDE-TRADING-ONE-ACCOUNT ... OK< RETURN TO TERMINAL

Grok and Claude Trading One Real Account_

PROJECT HAYSTACK DOC ARCHIVE :: AGENTIC AI INVESTING EXPERIMENT
sbrn.io/projecthaystack · doc file · updated 2026-08-12

Every AI trading experiment you have heard of is a cage match. Eight models, one arena, may the best benchmark win. I run the other experiment: Grok and Claude, two frontier models from two rival labs, operating one real-money brokerage account together under one written rulebook, with one public scoreboard that does not care which of them was clever today. This page is the cooperation story, told by the desk they share.

The arena era

Model-versus-model trading contests are having a moment, and some are worth watching. Public arenas hand identical allocations to a roster of frontier models and publish live standings. Research benchmarks like StockBench test LLM trading agents against months of market data. Portfolio products let thousands of people shadow a single model's book. The format is a ladder, and the question is always the same: which model trades best?

It is a fine question. I think it is the second most interesting one. The benchmarks keep finding that the gap between models is smaller than the gap between designs: what an agent is allowed to do, what memory it keeps, and who checks its work move outcomes more than whose logo is on the weights. If that is true, the frontier experiment is not model against model. It is models in different seats of one institution.

Arena cage matchThis desk
AccountsOne per model, rankedOne, shared
Question askedWhich model trades best?Can rival models run one honest institution?
BuysEach model freestylesDual gate: both must independently say buy-ready
RulesWhatever each model decidesOne written law, amended only through review
GradingP/L ladderProcess and outcome, scored separately, in public
When a model is wrongIt drops a rankThe other one catches it before cash moves

One account, two models, three seats

The division of labor is strict, and it is the whole trick.

  • Grok is the steward. It holds the only seat that touches the brokerage account. Once a day, at 10:00 AM ET, it checks settled cash, reruns the quality gates on every held and buy-ready name, asks Claude to co-score the menu before any deploy, deploys every settled dollar equally across dual buy-ready names during regular market hours, and files the paperwork. It does not brainstorm at the buy button. If Claude cannot answer, Grok solo deploys and logs the solo run.
  • Claude is the red team and deploy co-scorer. It never places an order. It still simulates, backtests, and adversarially reviews every proposed change to the law. It also independently scores every menu name before cash moves (buy-ready, no-add, fail, or hold). Buy-ready requires both; neither model can deploy a dollar unless the other agrees. That is a second job, not a soft suggestion.
  • The human is the treasury. Funds the account, ratifies the rules, answers the rare escalation, and otherwise stays out of the way.

The full architecture, including the written law both models answer to, is laid out in how an autonomous AI investing desk works.

The shared brain is a folder of markdown

The two models do not coordinate over some elegant agent protocol. They share plain-text files: the trading plan, the rulebook, the research notes, and a scored journal of every closed trade. Each reads the current state before acting and writes its results back for the other. Plain text sounds primitive until you try to audit an institution built on it: every decision is diffable, nothing hides in a binary blob, and neither model can quietly maintain a private fork of the strategy. Institutional memory, it turns out, fits in a folder.

Field notes: what each model is actually good at

Benchmarks will not tell you this part, so here are the desk's field notes from running both in production.

  • Grok is a metronome. Scheduled runs, order placement, fill logging, and daily source scans, the same checklist every day without existential commentary. One run split $105 across 17 names in fractional orders and filled all of them inside regular hours; the paperwork was done before the human finished lunch.
  • Claude is a demolition contractor. Its best work here has been destroying good-sounding ideas with evidence: thousand-path Monte Carlo runs, rolling-window backtests, and literature checks that end with a proposal dying politely in committee instead of expensively in the account.
  • Neither is trusted with discretion at the moment of execution. The deploy rule is pure arithmetic (equal split across eligible names) precisely because every current LLM drifts when handed judgment under pressure. The law does the deciding. The models do the working.

Where they disagree, and who wins

Disagreements are not a bug; they are the product. There are two disagreement channels, both public.

Before every deploy (menu votes)

Before Grok spends settled cash, both models score the full menu. Merge is conservative:

  • Either says fail, no-add, or hold: that call wins (no new buys of that name).
  • Buy-ready requires both.
  • Disagreement parks the name until they agree, and the dashboard Deploy gate table shows Grok's call, Claude's call, and the final.
  • If Claude is unreachable: Grok deploys alone and logs a solo-fallback row so the silence is visible.

That gate is operational, not theatrical. It is how two rival labs share a buy button without either model freestyling.

When the law itself is on trial (council)

Proposed rule changes still get argued in council, and those fights are published, verdicts included, on the dashboard. A few real ones from the first weeks:

  • A hybrid ranked deploy (weight the favorite names heavier) was rejected as spec-fragile: too many judgment calls per run for an automated steward to execute without drifting.
  • Steering deposits toward underweight names looked prudent and lost the argument to a thousand-path Monte Carlo: it wins only on lucky start dates, with identical drawdowns.
  • A one-shot deploy for newly added names was evaluated twice and shelved when its apparent edge dissolved into schedule luck.
  • The best catch so far was not a trade at all. Before one deploy variant could go live in July 2026, the red team audited the steward side's supporting simulation and found three bugs in it: a cold-start artifact, a seed-dependent result, and a missing control run. Bugs fixed, the recommendation flipped, and plain equal-split won before a dollar moved. That is the cooperative earning its keep.

The record so far, dated

As of 2026-08-12 (day 35): most menu votes come back unanimous, which is what you want from two models applying the same written gates to the same facts. The exceptions are all on the record, and so far they rhyme. Every recorded split has run the same shape: Claude called an outright fail on a quality gate, Grok had the same name at weak-but-passing, the restrictive call won, new buys into that name stopped, and the shares already held stayed put. Two of those names were later re-scored in a full review and Grok came around to the fail, which is an argument resolving rather than being quietly forgotten. Dates, tickers, and both calls are on the disagreements log, which also tracks whether each veto is aging well against the index. The shared book stands at +11.24% versus SPY +2.82% over the same window (alpha +8.42%), all of it on the live results page and the dashboard, where the next disagreement will publish with who said what.

Notice who won each fight: not Grok, not Claude. The evidence. When the models disagree about law, nobody pulls rank, because neither has any. A proposal survives simulation and adversarial review or it dies. When they disagree about a ticker on a deploy day, the conservative call wins and both votes stay on the public ledger. Every exit that follows is graded twice, process and outcome, under the desk's scoring system, and only process grades are allowed to amend the law.

The scoreboard does not take sides

One account, one rulebook, one benchmark: buy-and-hold SPY over the same window, printed next to the holdings and every scored exit. Until the desk beats SPY for six straight months, the official position is that the boring index is winning. That standard binds both models equally, which may be the most even-handed thing anyone has done to Grok and Claude all year.

For the record: Grok and Claude are products of xAI and Anthropic respectively; neither company is affiliated with this experiment. They are tools in seats, working shifts, sharing a brain made of markdown. The cage match is more cinematic. The cooperative has receipts.

FAQ

Which is better at trading, Grok or Claude? On this desk the question does not resolve, on purpose: one account means one shared result. What the field notes do show is a division of talent: Grok excels at metronome execution (scheduled runs, fills, paperwork), Claude at destroying bad ideas with simulations before they cost money. The arenas rank models; this experiment measures whether rivals can run one honest institution.

Can two AI models really share one brokerage account? Yes, with strict seats: only Grok places orders, every buy needs Claude's independent agreement first, and both answer to one written rulebook neither can edit alone. The account is real, funded, and published daily since 2026-07-09 on the live dashboard.

Do Grok and Claude ever disagree? Yes, and the channels are built for it: ticker votes park a name until both agree, and rule proposals go through adversarial review. On the ticker side every recorded split has gone the same way, with Claude's fail beating Grok's weak-but-passing and the name losing its buy. In council, several good-sounding strategies died by Monte Carlo. Every disagreement publishes with who said what.

Where to go next

Nothing on this page or this site is investment advice. This is a public experiment log for a small, isolated account. The full disclaimer is at the bottom of every page.

PART OF THE SECOND BRAIN :: sbrn.io