Rule Adherence Score — Grade Process, Not P&L
1 / Why this question showed up
A trade can make money and still be a bad decision. Another can lose money while following the written plan exactly. If a journal quietly grades decisions by outcome, random reinforcement teaches the opposite lesson: keep the lucky improvisation, drop the sound rule after a normal stop-out.
So the question is narrower than “was I profitable?” It is: did this decision match the rules that were declared before the trade, using only information available at the time?
That separation is what a rule adherence score is for — a process mark you can review beside P&L, then connect to results only across a sample large enough to mean something.
2 / What “rule adherence” means here
Rule adherence is the share (or weighted degree) to which a trade follows predeclared entry, risk, management, exit, and state rules. The rules have to be observable: you — or another reviewer — should reach the same pass/fail from the same record.
| Layer | What you check |
|---|---|
| Entry | Approved setup, confirmation, session, and timing were present. |
| Risk | Stop, size, and total exposure stayed inside limits. |
| Management | Partials, trailing rules, and no-add conditions were followed. |
| Exit | Planned target, invalidation, time exit, or declared discretionary protocol. |
| State | Daily loss, news, fatigue, and no-trade restrictions respected. |
| Documentation | Required plan note / screenshot captured at the defined time. |
“Take only good setups” cannot be scored. “Long only after the 15-minute close above the opening range, before 11:00, stop below the trigger low” can. If a score keeps causing debate, fix the written rule before chasing a higher percentage.
3 / Start with a minimum viable scorecard
The useful scorecard is short enough to finish on every trade. Begin with four to six rules that protect edge or account risk — not a duplicate of every sentence in the plan, and not cosmetic preferences scored as if they equal a hard risk limit.
- Write one observable pass condition per critical rule.
- State when evidence must be captured (before entry, during management, after exit).
- Define whether a violation makes the whole trade off-plan.
- Test the wording on a handful of historical examples for ambiguity.
- Freeze the wording for a review cycle before comparing scores across weeks.
4 / Binary model, graded model, critical fails
Binary: pass or fail
Mark each applicable rule 1 if followed, 0 if violated. Non-applicable rules are N/A — not automatic passes.
Arithmetic illustration only: if five of six applicable rules pass, the score is 5 ÷ 6 × 100 ≈ 83.3%. That is definitional math, not a claim about any book.
Binary scoring is fast and repeatable for clear constraints (“risk ≤ 0.5%”, “no entry after 11:00”, “never widen the stop”). Its weakness is severity: missing a screenshot and doubling risk both score zero unless a critical-fail rule distinguishes them.
Critical fails without hiding the component score
Predeclare a small set of safety violations that make the trade off-plan regardless of the percentage — for example exceeding maximum risk, trading after a daily stop, removing the protective stop, or taking an unapproved setup. Keep the numerical score as detail; classify the trade off-plan. Do not invent a “critical” rule after a loss.
Graded / weighted (only where partial means something)
Some decisions are not naturally binary (documentation complete / partial / absent; exit discretion inside a defined band). A graded model can assign 0, 0.5, or 1 if each level has an observable definition.
- Prefer binary when rules are objective and speed matters.
- Use graded scoring only where partial compliance is well defined.
- Use weights sparingly and write down why one rule matters more.
- Keep critical risk limits as hard fails even in a graded model.
- Do not retune weights to make a disappointing month look disciplined.
5 / Trade score vs period score
A week’s adherence should aggregate applicable rule decisions — not merely average trade percentages when trades have different numbers of applicable rules.
Also report the share of trades that were fully on-plan. Component adherence and fully-on-plan share answer different questions: frequency of rule hits versus complete execution. One can be high while the other is mediocre if small violations are spread across many entries.
6 / Compare on-plan and off-plan groups — carefully
Once trades are classified, performance can be reviewed separately. Expectancy in R keeps different cash sizes comparable:
or equivalently: total realized R ÷ number of trades
Compute that independently for on-plan and off-plan groups when you have enough trades to bother. Include costs and slippage in realized R. Report count beside expectancy so a three-trade subgroup is not treated as established.
Related reading on how people summarize results (without turning this memo into a metrics lecture): Performance metrics.
7 / Do not rewrite the strategy because off-plan trades lost
An off-plan loss is weak evidence about the planned strategy — the strategy was not executed. If a breakout rule requires a close above resistance but the trader anticipates the close and loses, changing the breakout stop or target from that loss uses contaminated evidence.
- On-plan trades test the strategy.
- Off-plan trades test the cost and triggers of deviation.
Fix execution from the second group; evaluate strategy rules primarily from the first. That does not mean ignoring the loss — record its R impact, review the trigger, install a process control. Strategy changes belong later, when on-plan evidence (and a stable rule version) actually supports them. Sample-size humility belongs here too; see the broader validation mindset in Validation gauntlet.
8 / Opportunity quality ≠ execution quality
Trade-only adherence can look perfect while valid setups are skipped. It can also punish restraint if every no-trade is scored as a miss. A fuller review starts from qualified opportunities: moments that met the plan whether or not a position was opened.
| Case | Reading |
|---|---|
| Valid & taken on-plan | Correct recognition and execution |
| Valid & skipped | Hesitation, capacity limit, or legitimate no-trade rule |
| Invalid & taken | Selection failure or impulse |
| Valid & taken off-plan | Real opportunity, broken execution |
| Invalid & skipped | Correct restraint (often logged only in aggregate) |
False-action rate = invalid trades taken ÷ all trades taken
Define the scan universe and hours first, or the opportunity denominator becomes arbitrary. Keep setup quality and execution quality in separate fields.
9 / Compact mistake taxonomy (and careful “cost”)
The adherence score says whether a rule failed. A small mistake category says what kind of control should prevent recurrence. Prefer one primary category plus one specific behavior (“Risk → moved-stop”) over a pile of emotional labels.
- Selection — entered without required setup/context
- Timing — early, late, or outside approved session
- Risk — oversize, correlated exposure, missing stop, daily limit breach
- Management — unauthorized add, moved stop, unplanned partial
- Exit — fear exit, ignored invalidation, missed time rule
- Process state — revenge, FOMO, fatigue, checklist bypass
- Documentation — missing evidence that blocks reliable review
Do not encode the outcome (“bad loss”, “lucky win”) as a process category. On cost: realized R of mistaken trades is an impact total, not automatically the causal cost — an on-plan version of the same trade might also have lost. When a compliant alternative can be reconstructed without hindsight (e.g. stop should have closed at −1R but was widened to −2.2R), the excess is roughly the execution gap. When it cannot, report impact without claiming a counterfactual winner. Risk framing that sits next to this habit: Surviving risk.
10 / A practical review cadence
- After each trade — score while memory is fresh; attach evidence; classify deviation. Do not redesign the strategy in the heat of the fill.
- Daily — critical failures, daily loss status, unlogged opportunities, stop/go. Account protection, not strategy conclusions.
- Weekly — component adherence, fully on-plan share, on-plan vs off-plan R (with counts), opportunity capture, highest-impact recurring mistake. Commit to one observable control for next week.
- Monthly / per sample — rolling adherence by stable strategy version; audit score consistency; merge duplicate mistake tags; ask whether a repeatedly failed rule is unclear, impractical, or ignored. Strategy edits belong here only when on-plan evidence supports them.
A 100% adherence week can still lose money. A 60% adherence week can still make money. Process and outcome can diverge sharply in short windows; relate them only over a broad enough sample — and do not lower the pass threshold after a bad week, add rules solely because one trade lost, or remove a valid rule because breaking it produced a winner.
11 / What this page holds vs what it does not claim
① Hold onto
- Process vs outcome as separate grades
- Observable rules; N/A for non-applicable checks
- Critical fails as predeclared hard offs
- Period aggregation by rule units, not naïve averaging
- On-plan evidence for strategy; off-plan for execution controls
- Opportunity log separate from execution score
② Not claimed here
- Any Quant for Free account adherence % or expectancy
- That a high score guarantees profit
- A preferred journal vendor or scoring UI
- Universal weights, thresholds, or “best” taxonomy
- Market timing, sizing, or “how to trade”
12 / Continue on the site
Score the decision; leave the story for later
No obligation to publish a time series. The point of revisiting is whether “on-plan / off-plan” stayed honest when money argued otherwise.
- Pick four to six observable rules and freeze wording for one review cycle
- After a few weeks, check whether two reviewers would agree on the same chart + plan
- If one rule fails repeatedly, ask clarity / capability / environment / impulse before rewriting the edge