ROBCO INDUSTRIES (TM) TERMLINK PROTOCOL :: NIGHT CITY INVESTING SUBNETACCESSING DOC ARCHIVE :: HOW-WE-SCORE-AI-TRADING-TRADES ... OK< RETURN TO TERMINAL

How We Score AI Trading Trades_

PROJECT HAYSTACK DOC ARCHIVE :: AGENTIC AI INVESTING EXPERIMENT
sbrn.io/projecthaystack · doc file · updated 2026-08-12

A green P/L line is a terrible teacher if you let it grade the homework. Lucky rule-breaking looks like skill. Disciplined losses look like failure. Over enough cycles, that inversion trains an AI desk (or a human one) to improvise.

I score every closed trade twice. The process grade asks whether the written law was followed. The outcome grade asks whether the trade made money. Only process is allowed to change the rules. Outcome stays on the report card so nobody pretends the market did not happen.

The live receipts sit on the dashboard scoreboard, and the running account numbers they feed live on the results page. This page is the rubric behind those rows. Methodology and memory live in the Second Brain: the shared plain-text plans, journals, research, and rules the agents read and write between runs.

Nothing here is investment advice. It is an experiment's grading policy.

Process vs outcome in trading: why one grade is not enough

Markets pay and tax randomly in the short run. Process is what you can control inside a rule set. If you merge the two into a single "good trade / bad trade" label, three distortions appear:

1. Rule-breaking winners get promoted. I learn that exceptions print money. 2. Rule-following losers get punished. I learn to abandon the plan during ordinary drawdowns. 3. The law becomes optional. Optional law is not law. It is a suggestion with a brokerage login.

Separating grades keeps the learning signal clean. A process A with an outcome loss is a good trade on a bad day. A process F with an outcome win is a problem that paid for its own cover story.

For how this loop sits inside the multi-agent desk (steward, red team, operator), see How this autonomous investing desk works.

Process grade rubric

The process grade answers one question: relative to the written law at the time of the trade, did I do the job?

Typical inputs:

  • Was the name eligible on the buy menu (buy-ready), correctly hold-only, or correctly NO-ADD after a filter fail?
  • Did the deploy follow equal-split rules and cash gates when the trade was an open (no new buys into NO-ADD, hold-only, or 20% capped names)?
  • Did the daily quality-gate recheck run, and were status flips (buy-ready / NO-ADD) recorded when filters changed?
  • Was the exit a hand-placed operator decision (the desk is buy-only and never sells; law #39), a ratified plan change, or an unauthorized improvisation? Any agent-placed sell is an automatic process F. Soft filter fail must not invent even a flag beyond NO-ADD.
  • Were hard bans respected (margin, shorts, short options, sub-365-day options, market-hours policy)?
  • Did the two-agent co-score gate run before cash moved (both models independently buy-ready, any disagreement honored), or was a solo fallback properly logged when the second model was unreachable? (Gate live since 2026-07-17.)
  • Was the journal complete enough that a later agent can audit the decision without folklore?

Letter scale used on the public ledger:

GradeMeaning
ALaw followed end to end. Documentation complete. No hidden exceptions.
BMostly followed, with a documented higher-level plan migration or minor process debt that is already recognized in the journal.
CMaterial process miss: incomplete checks, ambiguous trigger, or sloppy paperwork that still landed near the rules.
D / FClear violation or unrecoverable process failure. Outcome is irrelevant to promotion.

Process grades are assigned from the contemporaneous rules and notes, not from post-exit price action. The market does not get a vote on whether the steward followed the schedule.

Outcome grade rubric

The outcome grade is deliberately boring.

LabelMeaning
WinRealized P/L greater than zero after fees and fills as recorded.
LossRealized P/L less than zero.
FlatEffectively zero (including intentional residual closes that neither helped nor hurt).

Optional context fields (R-multiple on premium, percent on equity, hold time) may appear in the journal for analysis. They do not override the win / loss / flat label, and they never rewrite the process grade.

What can change the rules

Only patterns in process grades, plus red-team review, can amend the law.

A streak of outcome wins under broken process is treated as a bug report, not a strategy upgrade. A streak of outcome losses under clean process may still trigger research (menu quality, ban list, deposit cadence), but "the market was mean" is not by itself a license to invent dip-buying logic at 10:00 AM.

Proposed rule changes are simulated and adversarially reviewed before they ship. Rejections stay public in the council review section on the live dashboard. The ledger backs this up: in the first ten days the council rejected six proposals (ranked deploys, underweight steering, one-shot deploys for new names, momentum and dip variants, sell overlays) and adopted three, all three of which reduce discretion rather than add it. The Second Brain keeps the append-only history so retired tactics cannot quietly return without a fight.

Worked examples (from the real ledger)

Record to date: three closed legs, two wins and one flat, +$39 realized, every leg graded process B, zero A grades and zero violations to grade yet. All three are Plan A retirements, my first plan being shut down, not the current strategy succeeding; the honest labels below are the point of the system.

Real row: plan retirement close (PATH call, +$19) A September $13 call on PATH, entered at $89, exited at $108, +$19. No take-profit fired; the position closed because the whole short-dated options plan was retired the same day the evidence said it could not beat a SPY drip. Process grade B: the higher-level migration was ratified and journaled, but the original hold-to-target plan was broken. Outcome: win. Lesson that shipped: plan pivots must journal as plan-retire and cancel working take-profit orders explicitly, so the stats never dress a pivot up as a target hit.

Real row: quick green on a lottery-ish strike (ZETA call, +$20) A ZETA call entered at $75 closed at +$20 in the same retirement. Same process B, same honest label: the win validates nothing about the strike selection, and the journal says so. A fast green on a speculative leg is exactly the kind of evidence this scoring system is built to refuse.

Real row: the flat residual (+$0) A $35 residual PATH leg closed flat when the plan retired. Process B, outcome flat. Lesson: filler tickets add noise without information, and that class of position is now banned in the live law (no short-dated options, no filler structures).

Example pattern: rule-breaking winner (still hypothetical, happily) Suppose an agent sold a name solely because the day was red, with no thesis death and no fraud trigger. Outcome could be a win if price fell further. Process is still a failure. The correct institutional response is write-up and, if needed, tighter automation, not a parade. The ledger has zero of these so far; the system exists so the first one gets caught.

You can inspect the actual rows, dates, and wording on the live closed trades scoreboard.

Anti-patterns this scoring system is designed to kill

1. Outcome worship. "It made money, so the process was fine." 2. Hindsight rule editing. Changing the law after a loss to pretend the loss was compliant. 3. Screenshot science. Publishing winners without the process column. 4. Silent exceptions. One-off "just this once" deploys that never hit the journal. 5. Confusing agent flexibility with permission. Agentic systems need sharper grades than fixed bots, because they can invent new paths. See Agentic AI trading vs trading bots.

Equal-split deploy is process-friendly on purpose: fewer judgment calls per run means fewer gray-zone grades. Background on that choice: Equal-split deploy explained.

How to read the public scoreboard

When you open the dashboard:

1. Read which rule fired before you read P/L. 2. Compare process and outcome as separate columns in your head even when both are short labels. 3. Prefer lessons that produce rule clarity over lessons that produce cleverness. 4. Check whether a rule change was proposed. Most good lessons do not need a rewrite. 5. Zoom out to SPY over the same window on the dashboard. Trade grades teach process. The benchmark teaches humility.

If you are still mapping the category, start with What is agentic AI trading?, then return here when a closed trade row looks sparse without context.

Where to go next

Nothing on this page or this site is investment advice. Scoring language describes how this experiment evaluates its own behavior. It is not a promise of future results, a recommendation to copy trades, or a claim that process grades produce profits. I run a small, isolated account. Do your own research and consult a licensed professional before making investment decisions.

PART OF THE SECOND BRAIN :: sbrn.io