Survivability · Capital

Why Trading Firms Die

The infrastructure and risk-ops failures that kill profitable strategies.

Three deaths, three speeds

Knight Capital was one of the largest market makers in US equities. It lost $440 million — and its independence — in forty-five minutes. Alameda Research was one of crypto’s largest market makers and OTC desks, running its own quant book — and FTX, its sibling exchange, was one of the largest on earth. From a leaked balance sheet to bankruptcy for both, an eight-billion-dollar hole and all: about a week. Situational Awareness was an AI-thesis hedge fund reportedly running $45 billion at its peak; a month of margin calls later, it was fire-selling its book to a rival.

Different eras, different markets, different asset classes — and one thing in common: not one of them ran out of alpha. Knight’s market making worked. Alameda’s desk and FTX’s exchange both printed money. The fund’s thesis may yet prove right. What failed was everything around the strategy — a deployment, a risk structure, a hedge that wasn’t one. The strategy is almost always fine.

Most trading firms don’t die because they run out of alpha. They die because something else breaks — and it breaks faster than they can recover. After eight years operating a crypto trading system across every regime since 2018, we’ve watched firms with better capital, better talent, and better signals than us disappear. Almost none were killed by a bad strategy.

The myth of better strategies

Ask a fund operator what keeps them up at night and you’ll usually hear a version of the same answer: our edge is decaying, we need better signals, we need faster models. Alpha generation gets most of the hiring budget, most of the intellectual energy, and nearly all of the ego.

But the graveyard tells a different story. The famous deaths — and the quiet ones nobody writes about — trace back to deployments, counterparties, reconciliation gaps, and concentration. Survivability is a function of three things: operations, risk management, and strategy diversification. Alpha decides how much you make. Those three decide whether you’re still here to make it.

The myth survives contact with the graveyard because it isn’t an information problem — it’s an identity problem. In this industry, alpha is where identity lives: a signal is proof of intelligence, the thing that makes a quant a quant. Nobody’s self-image — or comp narrative — is built on their kill switch. So the smartest people compete to work on signals, and budget follows self-image rather than failure probability.

Most firms overinvest in alpha and underinvest in the other three. They build a Ferrari engine and bolt it to a go-kart frame.

Why this business dies differently

Trading firms increasingly look like tech companies. They can’t operate like them. Four structural differences explain why:

Capacity is endogenous. In SaaS, more customers means more revenue. In trading, more capital doesn’t mean more returns — every strategy has a point at which your own orders move the market against you. Scaling means finding new edges, not adding capital to old ones.

Mistakes compound in milliseconds. Most businesses can ship fast and fix later. In trading, the feedback loop between error and consequence is measured in milliseconds, not sprints. A bug that runs for 45 minutes can end the firm.

A/B tests cost P&L — and the arms interact. We run live A/B tests: parallel variants of a strategy trading real money side by side. It works, but it is nothing like shipping a feature to 5% of users. Every arm pays for its own sample in realized P&L, and unless the experiment is set up carefully the variants self-compete — queueing at the same prices, splitting the same fills — or quietly confound each other’s results. In software the control group is free; in trading there is no free control group. Backtests and paper simulation help, but the gap between simulation and live is where fortunes are lost.

The operating model is a hybrid. A trading firm has to run as three organizations at once: a technology company ( complex software, continuous deployment), a financial institution (risk, capital, counterparties), and an operations center (real-time monitoring, incident response, zero tolerance for downtime). Most founders deeply understand one of these and have a passing familiarity with the other two. The gap is where failures live.

Put those together and you get three distinct ways to die, sorted by speed:

  • Death by minutes — an operational failure outruns your ability to stop it.
  • Death by months — a risk or counterparty exposure carried unpriced for months, until a trigger arrives and the ending takes days.
  • Death by years — concentration in an edge that quietly stops working.

Every trading-firm failure we know of fits one of these three. The next section takes them one at a time.

Three deaths

Death by minutes — operations

August 1, 2012. Knight Capital — then one of the largest market makers in US equities — was going live on a new NYSE retail program, with a launch date fixed by the exchange and roughly a month to build for it. The deployment was manual: in the days before launch, a technician copied the new code to the production servers by hand, with no second engineer reviewing the copy and no automated check that it landed everywhere. It missed one of the eight. That eighth server still carried a dormant module from years earlier, and the new code had reused its old activation flag — so at the open, the orphaned server began firing orders from a defunct strategy at full speed. Then, under pressure to fix it live, engineers diagnosed the new code as the culprit and rolled it back from the seven healthy servers — spreading the defect to all eight. Forty-five minutes and several billion dollars of unwanted positions later, Knight had realized a ~$440 million loss — roughly four times its prior-year earnings. It survived only by emergency rescue financing and was absorbed by a competitor within the year.

The lesson isn’t the specific bug — it’s that nothing in the chain was a trading decision. A hand-run deploy under a deadline someone else set, no verification that it landed, and a panicked fix that made it worse: process failures, priced by the market at $440 million. The thing that kills you won’t be the thing you’re optimizing for. It will be the thing you assumed was fine.

None of what killed Knight requires heroics today — the fix is the same automated deploy discipline any serious software shop runs, and the one our own pipeline enforces: a release either passes its gates — interface compatibility, dependency and version checks, tests — or it never reaches production; what shipped is verified to be what’s running on every node; and rollback is a tested path, not a 4 a.m. improvisation. The rest of the controls that prevent death by minutes are just as unglamorous: kill switches that don’t ask permission, position reconciliation that runs continuously rather than nightly, and alerting that reaches a human who can act. None of them generate a basis point of alpha. All of them are why you still have a book.

Death by months — risk and counterparty

FTX wasn’t a trading failure; it was a risk-isolation failure inside an otherwise dominant business — client funds commingled with its trading arm, Alameda Research, an exposure that months of massive revenue could not paper over. We were trading on FTX when it collapsed in November 2022. Our risk systems flagged abnormal withdrawal behavior, we triggered our exit protocols, and we got the large majority of our funds out before withdrawals froze. That week gets its own piece — Same Starting Line, Different Paths — but the short version: counterparty risk is not a footnote in crypto; it is a first-class risk, and it nearly cost us the firm’s capital while we watched others lose everything.

And this death is not crypto-specific. In July 2026, Situational Awareness — Leopold Aschenbrenner’s AI-thesis hedge fund, reportedly running roughly $45 billion at its peak on leverage of up to 400% — was forced by prime-broker margin calls into a fire sale of its leveraged public positions after both sides of its book moved against it at once: the long AI-infrastructure bets fell while the bearish software positions meant to hedge them went the wrong way. Reported assets dropped from about $45 billion to around $10 billion inside a month. The thesis may yet prove right. The leverage — and the hedge that wasn’t a hedge — decided the outcome before the thesis got a vote.

Notice that both endings were fast — FTX went from a rival’s tweet to a withdrawal run to bankruptcy in about a week; Situational Awareness’s fire sale played out in days once the margin calls started. That’s because this death runs on two clocks. The exposure builds quietly for months or years — a venue trusted too much, a bank, a bridge, a margin arrangement, a hedge that only works in calm — carried every day without being priced. Then a trigger lands: a tweet, a leaked balance sheet, a margin call. The collapse that follows is measured in days, but its speed is a property of the trigger, not the cause. You die in days; you were killed over months. The death is named for the slow clock because the slow clock is the only one you can manage — nobody has ever risk-managed a bank run from inside it.

The controls here are structural, not clever: exposure limits per venue and per counterparty, capital that moves on triggers rather than opinions, isolation between books, and reconciliation that treats every external balance as a claim to verify, not a fact.

Death by years — concentration and decay

The slowest death gets the least attention because there’s no headline — just a strategy whose edge erodes until the firm quietly returns capital. Edges get crowded. Market structure shifts. Fee schedules change. Regimes end. Every strategy has a half-life, and a firm concentrated in one edge is a firm with an expiration date it hasn’t computed.

We know this one from the inside, because we’ve killed our own strategies. Over eight years we’ve retired entire strategy classes after the data said the edge wasn’t there — including approaches that looked excellent in backtests and that much of the market still actively trades. That map of what doesn’t work cost us real money and real years to draw, and it’s why we treat diversification as infrastructure: the system that detects decay and rotates capital is worth more than any single signal it monitors.

The pattern repeats across whole strategy classes. Simple cross-venue arbitrage printed money in crypto’s early years; those spreads are now competed down to the cost floor of the fastest, cheapest operators. Passive market-making on the major pairs went the same way — spreads compressed until adverse selection swamps the rebate for anyone without a structural advantage. Simple mean reversion had a long profitable run before crowding ground it toward zero. None of these edges “failed.” They matured: the returns migrated to the handful of firms with the lowest costs and the fastest infrastructure, and much of the market is still paying tuition to rediscover that.

The control against death by years is a portfolio and a process: multiple uncorrelated edges, continuous performance monitoring against expectation, and the discipline to kill a strategy the data has already killed — before it takes the firm with it.

Our own entries in the ledger

Every firm that lasts writes entries of its own into this ledger. Ours are small so far — partly by design, partly by luck — and we keep one from each speed.

By minutes. The entry we study hardest is a death that didn’t finish. On an otherwise ordinary evening, our hourly state verification came back red: the book our strategies believed — flat, orders closed, nothing at risk — did not match the venue’s. A venue had confirmed orders as filled without delivering the records of the fills themselves, and the fallback built for exactly that gap had been disabled by the same code path that hit the bug. The check that caught the drift then hit a dead end: an error state with no escalation path, and the kill switch — the component built for exactly this moment — crashing and retrying in a loop, unable to finish the one job it exists for. We reconciled the books by hand and restarted. Within minutes the same mechanism dropped dozens more fills — the restart had fixed the data, not the mechanism — and strategies stacked fresh positions on top of exposure they couldn’t see, until the venue’s own balance limits stopped them. The ledger says we made a small profit that night: the unintended exposure — a few thousand dollars of notional, at our deliberately small scale — happened to close into a rally. We recorded it as a failure anyway; the same mechanism with the price moving the other way is a loss with no bound we controlled. What came out of it: a fill pipeline whose fallback the primary path cannot disable, an intermediate order state that blocks new risk while fills are being recovered, automatic strategy pause on repeated rejections, a kill switch that distinguishes fatal from transient errors — and a restart rule that requires proving the pipeline healthy, not just the books balanced.

By months. The FTX collapse handed us a controlled experiment in what counterparty risk actually is. One of our withdrawals showed “processed” on FTX’s side, but the transfer was never made — those funds never left the exchange, and they came back only years later, through the bankruptcy queue. Our bank at the time was Silvergate, the crypto industry’s banking backbone and the next domino in the same blast radius: within months it was gone too — and it returned every dollar we held there. Two failed counterparties, weeks apart, opposite outcomes. The difference wasn’t luck; it was structure — regulated, segregated deposits wound down in order, while commingled exchange balances became claims in a line. That’s the lesson we priced in: counterparty risk is not the size of the exposure but the structure holding it. A dollar on an exchange, a dollar at a bank, and a dollar in transit are three different risks — and each place money rests gets its own limit.

By years. The slow-death entry is passive market-making on the major pairs. The thesis was the oldest one in the book: quote both sides, earn the spread, collect the maker rebate — the spread as income. Seven versions of signals and gates tried to make it true, and each backtest found a reason for hope. What ended it wasn’t a better version; it was a diagnosis — sitting down with the fill data and measuring what each live fill was actually worth. The answer: the price moved against us after nearly every fill, by more than the rebate paid, and the damage didn’t shrink no matter how long or briefly our quotes rested in the book. That last fact killed the comfortable story that speed would fix it. This wasn’t a race we were losing; it was a seat at a table where our fills were the ones informed flow wanted us to have. So we killed the class instead of shipping the next version — the conclusion was structural, not a tuning gap: at our fee tier and scale, on those symbols, the spread was nobody’s edge. That sentence cost months of engineering and real trading losses to earn. It costs nothing to obey — and it’s on our map in ink.

That is what the three deaths look like from the inside, at survivable scale: an unwind caught in minutes, a counterparty lesson that cost patience instead of principal, a strategy killed by data instead of drawdown. The goal isn’t avoiding entries in the ledger. It’s making sure every entry is small enough to learn from.

What survival actually costs

Here’s the uncomfortable arithmetic. The three pillars — operations, risk, diversification — are not features you add to a trading system. They are the system. A realistic build for a serious crypto trading operation covers market data and connectivity across venues, an execution layer, a risk engine with real limits, backtest and validation infrastructure, monitoring and incident response, and the reconciliation plumbing underneath all of it. Built from scratch, that’s a 12–24 month effort by a team you have to hire first — in a market where the engineers and quants who have done it before are scarce, expensive, and mostly already employed by firms that pay them to stay.

And the clock runs at both ends. Every month spent building is a month of fixed burn with zero revenue — capital raised, mandate signed, edge waiting, nothing live. Time-to-live is usually the largest single line item in a new desk’s real cost structure, and it never appears in the budget.

That’s the trap: the pillars that keep you alive are exactly the ones that feel optional when you’re racing to deploy. Firms cut them to get live faster, and then join one of the three graveyards above.

Where this leaves you

If you allocate capital: the diagnostic is short. Ask a manager how they’d survive each death. What stops a runaway algorithm, and has it ever fired? What happens to the book if their largest venue freezes withdrawals tomorrow? How many uncorrelated edges are they running, and when did they last kill one? A manager with real answers has a durable operation. A manager who talks only about their signal is showing you the go-kart frame. If you’re evaluating crypto-native quant exposure, that’s the conversation we’re set up to have.

If you’re standing up a desk: the same three questions apply to your own build, on day one — not after the first incident. If the honest answer is that operations, risk, and validation are “phase two,” that’s the gap to close before the first incident, not after — build it, hire it, or partner for it. How matters less than when.

Either way: we’re writing this series from a live book, not a whiteboard. Which of the three deaths is your operation least covered against — minutes, months, or years? We read every reply, and the follow-up pieces will dig into the answers we hear most.


Martian Mobile is a proprietary crypto trading firm operating since 2018. This is the first piece in a series on survivability, building trading operations, and validating edges — the full series lives here.