← Back to blog

Bug Patterns · Live Bot Audits

The Risk Gate That Never Gets Its Input

A loss limit can be perfectly written and still never fire — because the number it compares against is zero, stale, or never delivered. Five variants of the same bug, including two from my own bot.

2026-10-11

When I review a trading bot's risk code, the first thing I look for is not whether the limit is correct. It's whether the limit ever receives anything. A daily-loss check that compares daily_pnl against a threshold is textbook-correct. If daily_pnl is never updated, it compares zero against the threshold forever and never trips. The gate is there, the tests for the gate pass, and the account is unprotected.

Over the last audits this has been the most consistent single pattern: the gate exists and is correct, and the signal feeding it is missing, frozen, or wrong. Five variants below. Three come from public repositories where I filed a normal issue (repositories and authors left unnamed); two are from my own bot.

Variant 1: the number that stays zero

A trading agent gated new orders on a daily loss limit. The gate read a running dailyPnl, and that figure was only ever updated by adding the PnL of a position at the moment it closed. The PnL field on every position was set to zero when the position was created, and nothing else in the file ever changed it. So every close added zero. A day of losses and a day of nothing looked identical to the check, and the limit was reachable in the configuration and unreachable in practice.

Variant 2: the breaker nobody calls

In another bot the main loop correctly checked a circuit-breaker flag and skipped trading when it was set. The flag could only be flipped by two paths. One was a method that records a trade's result, and in the whole tree that method was called only from unit tests. The other, a drawdown check, had no call sites at all, not even in tests. The gate was well built, and the thing that should have set it off was never connected.

Variant 3: the risk manager that never sees a sale

A third bot had two parallel loss-streak mechanisms. One had a method that computed the realized PnL of a sale and paused trading after a streak of losses; it was defined and never called. The other was called only from the buy path, where the profit-and-loss argument is always zero because nothing has been sold yet. Meanwhile the code that actually closed positions computed the real realized PnL correctly, and used it for a notification and for analytics, but never handed it to either risk mechanism. The one number the breaker needed was calculated and then sent elsewhere.

Variant 4: my own bot — the calculation that crashed on its first row

This one was in my own bot, and I only found it by watching a long paper run. The position sync read an attribute with a typo (avgCost instead of the library's averageCost) and raised on every call that included a real portfolio. The whole loop sat inside one try/except that logged the error and moved on. Result: the portfolio-exposure figure showed 0.0% for the entire session, not because exposure was zero, but because the calculation died on the first line of the first portfolio row, every time. The exposure limit existed, was correct, and had never once seen a real number. There happened to be no open positions, so nothing bad happened — luck, not design.

Variant 5: my own bot — reading the result from before it existed

My bot's loss-streak cooldown, which escalates the pause after consecutive losing trades, read each trade's PnL from the decision object at the moment of entering the trade. A trade that hasn't been closed has no PnL yet, so the value was always the default zero and the streak counter could never grow. The escalation logic was correct. It was just reading a field before the data it needed could exist; the fix moved the update to the handler that receives the real realized PnL from the closing fill.

This is one failure shape out of several I check for. The full checklist — entry-price path and exit/risk-limit path together — is on one page, or I can walk through it against your bot directly.

Get the safety-net checklist →

Why the tests don't catch it

A unit test for a risk gate sets the input itself: "given a daily PnL of -$500, the limit trips." That passes. It says nothing about whether the live system ever produces that input. The bug lives in the wiring between the gate and the rest of the bot, and wiring is exactly what a gate-in-isolation test skips. Variants 1, 2 and 3 all had a correct gate and a green test for it.

What to check in your own bot


The thread running through this and my other write-ups is that a safety net is judged by the event it has never been tested against. Earlier cases covered risk state that only lives in RAM, a guard that counts the wrong thing and a halt that a restart undoes. This one is the quietest: nothing is broken that you can see, because the check runs fine, on a number that was never there.

Contact

Get in touch

Not investment advice. Honest Backtest provides technical code review and backtest recalculation only — an honest read on what the code and the numbers actually show, not a recommendation to trade.
Photography via Unsplash.