Buying one options contract cost $347 in hidden fees. The same trade, one strike away, cost $7.
Nobody had to tell that trader. Stock trades come with a legally required receipt for this. Options don't — and options are where the money goes.
There's a price to buy and a lower price to sell, at the same moment. The gap between them is the house's cut. You pay it going in, and again coming out.
Open your broker, find the contract you are about to buy, and read off its bid and ask. Type them here. Nothing is sent anywhere — this runs entirely in your browser, and it grades your contract against the 65 we measured on live markets.
The gap is not a fee somebody set. It is the sum of three separate charges, and knowing which is which tells you when the gap is fair and when you are being quoted a price nobody expects you to take.
Somebody has to be standing there, willing to trade with you at 2:14pm on a Tuesday, in a contract that might trade eleven times all day. That readiness costs money — staff, exchange fees, capital — and the spread pays for it. In a competitive market this part shrinks to roughly what the business costs to run.
When they buy from you they are stuck holding it until someone turns up wanting the other side. Every minute of that is risk. The less often a contract trades, the longer that wait, and the wider the quote. This is why the illiquid strike one row down costs so much more than the busy one — nothing about your trade changed, only how long they expect to be stuck with it.
Some of the people hitting that quote know something. The dealer cannot tell which ones, so they charge everybody a premium to cover what they lose to the informed few. This is the uncomfortable part: you pay that premium even though you are almost certainly not the informed one. It is the price of being in the same queue as them.
Every bar below is a real contract we bought and sold on the same afternoon, on the same underlying, to find out what it truly cost. Sorted cheapest to dearest.
A deep in-the-money call moves almost exactly like 100 shares. Traders use them as a substitute. Here is what the substitution costs.
These are the same trade. A put spread where you sell the higher strike and buy the lower one, taking cash in, has an identical payoff at expiry to the call spread on the same two strikes that you pay to put on. Same maximum win, same maximum loss, same breakeven, same underlying price at every point. Only the direction the cash moves on day one differs, and put–call parity forces the two to cost the same once you account for it.
Getting paid up front feels like income. It is not income; it is a loan against a liability you have just taken on. The reason we still use the credit form is unglamorous: it puts the wider, busier strike on the side we are selling, which is where the liquidity is.
A limit order sitting in the book is a free option you have written and given to the market. You have promised to trade at your price; anyone may take that promise whenever it suits them. When the news is good for the other side, they take it. When it is bad for them, they leave it there. Your only protection is cancelling faster than they can act — and against a machine, you cannot.
This is why the professional advice is the opposite of the folk advice: use limit orders sparingly, priced close to the market, mainly where your size would otherwise move the price. Our agent crosses the spread deliberately and prices the crossing, rather than resting far away and pretending patience is free.
Selling an option pays you cash up front. It feels like free money. It isn't — the cash is priced against how likely you are to lose. Here's the honest arithmetic.
maxLoss / (maxLoss + credit). Widening the spread does not improve it — we checked at
1, 2, 3 and 4 dollars wide and all sit within half a point of each other. That is the market pricing
correctly, not a flaw in the trade.
The free data feed is an approximation. So instead of trusting it, we sent six orders at the same instant, each offering a wildly different price, and watched what we were actually charged.
| We offered up to | We actually paid | Difference |
|---|
The received wisdom is that patient traders get better prices. We ran it properly: ten pairs of identical trades, differing only in when the order went in. Each square is one pair.
You can't see the real price before you trade. But the free feed's quoted gap, unreliable as its exact value is, ranks contracts almost perfectly. Drag the line and watch the expensive ones fall away.
Everything above is deliberately plain. Everything below is not. This is the working: the models, the statistics, the regulatory basis, and the one finding we published and then had to withdraw.
Muravyev & Pearson (RFS 2020) argue option fair value is predictable at high frequency from the underlying. We tested it on 240 read-only samples across eight contracts, regressing each option's realised 40-second price change on the delta-gamma prediction δ·dS + ½γ·dS².
| Contract | Slope | t-stat | R² |
|---|---|---|---|
| SPY 0DTE ATM | 0.890 | 24.05 | 0.726 |
| SPY 1-day ATM | 0.727 | 17.48 | 0.584 |
| SPY 8-day ATM | 0.884 | 13.27 | 0.447 |
| QQQ 0DTE ATM | 0.645 | 14.40 | 0.487 |
| SPY 8-day OTM | 1.046 | 5.40 | 0.118 |
Every slope positive, every t-statistic between 5.4 and 24.1. The underlying explains 73% of the variance in the SPY 0DTE contract's 40-second move. Slopes below 1 are expected rather than a defect: the quoted midpoint updates in discrete ticks and so understates true moves, and implied volatility typically falls as the index rises, damping call moves.
| Horizon | SPY 0DTE | SPY 1-day | QQQ 0DTE | Signal ÷ feed noise |
|---|---|---|---|---|
| 2 s | −4.3% | −5.1% | −8.7% | 0.20 |
| 10 s | +10.6% | +4.5% | −5.9% | 0.69 |
| 20 s | +28.3% | +16.6% | +4.5% | 1.10 ← crossover |
| 40 s | +48.9% | +32.0% | +22.7% | 1.68 |
RMSE improvement over a random walk. The signal only clears the free feed's own noise floor at about twenty seconds — below that you are amplifying quote jitter, which is why a naive fast polling loop makes the estimate worse. Both predictors are scored against a noisy target, so common noise inflates both and pushes the ratio toward one: these are lower bounds.
Three design parameters fall out, none of them guessed: decision cadence of 40–60 seconds, SPY over QQQ, and stay near the money. Also worth noting — Alpaca returns no greeks whatsoever for same-day expiries, so the strongest row in the table is computed with our own Black-Scholes engine on an intraday clock. The platform gives you nothing exactly where the signal is strongest.
A null result is only interesting if the experiment could have detected the effect. Minimum detectable effect at 80% power, from the observed dispersion:
| Pairs | Smallest detectable effect | Detects the published ~$2.50? |
|---|---|---|
| 4 | $2.20 / contract | yes |
| 10 ← ours | $1.39 / contract | yes, comfortably |
| 77 | $0.50 / contract | — |
| 1000 | $0.14 / contract | — |
We ran ten pairs; the effect we were testing for needed four. The 95% interval [−$0.82, +$1.12] excludes it. We cannot rule out effects below about 1.4¢ a share — and we say so rather than claiming a cleaner result than we have.
And we can say why, not merely that. The waiting arm's raw fills came in 2.4¢ higher while the underlying rose 2.8¢ during the wait — yet the costs differed by 0.15¢. The real quote tracked the underlying almost exactly. Had it been stale, waiting would have paid. It didn't. Muravyev & Pearson sampled 2003–2006; market makers now requote in microseconds.
| Free feature | Run 1 (n=17) | Run 2 (n=48) | Verdict |
|---|---|---|---|
| Quoted width | +0.944 | +0.893 | robust |
| Option price | +0.873 | +0.789 | robust |
| Day volume | −0.735 | −0.416 | weakens |
| Ask size (depth) | −0.897 | −0.324 | did not replicate |
| Open interest | −0.500 | −0.282 | weak in both |
| Days to expiry | −0.105 | +0.413 | unstable |
Spearman rank correlation with true spread. On the first sample we reported ask size as the second-best predictor and noted that essentially no retail tool uses it. Replication killed it. It was an artifact of a small sample dominated by deep in-the-money contracts. Only quoted width and option price survive. The retraction stays on the page on purpose.
One folklore casualty worth noting: open interest is the weakest usable predictor and errs both ways. One contract with open interest of 1,703 — comfortably "safe" by the standard retail screen — cost $306. Another with 199 cost $1.
The Market Access Rule requires a broker-dealer to maintain pre-trade controls. Our fourteen gates map onto its clauses:
| Clause | Requirement | Our gates |
|---|---|---|
| (c)(1)(i) | Pre-set credit / capital thresholds | per-trade 0.25%, daily 1.0%, drawdown 2.5%, aggregate 1.5% |
| (c)(1)(ii) | Erroneous orders — price, size, duplication, rate | fat finger, price collar, liquidity gate, duplicate detection, throttle |
| (c)(1)(iii) | Regulatory compliance before entry | uncovered shorts unrepresentable, expiry cutoffs, account guard |
| (c)(2) | Immediate post-trade surveillance | every decision journalled with its reason |
The gates also have a second lineage, which is the list of ways derivatives desks have actually destroyed themselves. Hull devotes a closing chapter to it, and the lessons map onto this system one for one:
| Lesson from the disasters | How it is enforced here |
|---|---|
| Define risk limits at board level, then convert them into limits on individuals | Every threshold lives in one config, serialised and hashed into the pre-registration before the first order |
| Take the limits seriously — the penalty for breaching a limit must be the same when the breach makes money | The kernel has no override path. Vetoes of trades that would have profited are journalled and reported alongside the rest |
| Do not assume you can outguess the market: one trader in sixteen is right four quarters running by chance | The agent claims no forecasting edge. The regime signal can only withhold permission, never size a bet |
| Do not blindly trust models — large profits from simple strategies usually mean the model is wrong | The fill oracle exists because we did not trust the simulator's prices, and went and measured them |
| Separate the front, middle and back office | The agent proposes, the kernel vetoes and can only veto, the hash-chained journal and broker reconciliation record — three separate pieces of code |
| Barings and Société Générale: an arbitrage mandate quietly became a directional bet | Structure is fixed in code. The model may abstain or shrink; it cannot choose strikes, expiries, or direction |
Hull, Options, Futures, and Other Derivatives, ch. 35. The separation-of-duties point is the one worth stressing: Leeson controlled both the front and back office at Barings, and Kerviel had worked in Société Générale's back office before becoming a trader. An agent that both places orders and writes its own record of them has the same structural flaw.
The clause that shaped the design is (c)(1)(ii). Its erroneous-order controls exist to catch your own system malfunctioning, not the market moving. Retail risk management is almost entirely about market risk; professional pre-trade risk assumes your own code is the most likely thing to be broken. Knight Capital lost $460m in 45 minutes in 2012 to deprecated code on one of eight servers, and the SEC charged them under this very rule — finding they had compiled an inventory of controls rather than considered malfunctions in their own order router. Their system also fired 97 warnings before the open that nobody acted on.
So one gate has no market-risk purpose at all: after enough consecutive refusals the kernel halts, on the reasoning that an agent being refused repeatedly is broken rather than unlucky. Fourteen gates, twenty-three tests, no network required to run them.
Claimed
Not claimed
A short-volatility book showing a beautiful number over a few days is the predicted output of the measurement error, not evidence of skill. Goetzmann, Ingersoll, Spiegel and Welch (RFS 2007) show a fund that sells out-of-the-money options and holds cash has, whenever those options expire worthless, a zero standard deviation and a positive excess return — and therefore an infinite Sharpe ratio.
So the rules were fixed before the first trade. Risk limits, strategy and three falsifiable predictions were committed and hashed in advance — including the prediction that profit and loss over the evaluation window will not be statistically distinguishable from zero. Commit bac24e3.
This is the calculation that decided what this project would try to be. Active management has a standard formula for how much skill a strategy can express — the fundamental law, IR ≈ IC · √BR: your information ratio is your skill per decision multiplied by the square root of the number of independent decisions you make. Breadth is the lever, and it is the one a short evaluation window takes away.
An agent trading one strategy on one underlying over five sessions makes on the order of five independent bets — not fifty, because the same regime call repeated daily is one bet made five times, not five bets. Even at a skill level most professionals never reach, five bets cannot produce a distinguishable result. And the standard error of an information ratio is roughly 1/√years: establishing merely top-quartile skill at conventional confidence takes about sixteen years of returns. A five-day P&L number is not weak evidence of skill. It is no evidence of skill.
So we moved the breadth to where it exists on this timescale. Measurement has breadth: 65 contracts probed for the liquidity gate, 40 paired quote comparisons for the feed audit, 240 samples for the fair-value study, 10 matched pairs for the timing test, 10,000 randomised cases against the risk kernel. Those are questions a week can actually answer. The trading result is reported honestly and claimed for nothing.
| Quantity | Value here | What it implies |
|---|---|---|
| Independent bets in the window (BR) | ~5 | √BR ≈ 2.2 |
| Signal-to-noise of the P&L, pre-registered | ≈0.11 | P&L indistinguishable from zero |
| Years to establish IR = 0.5 at t = 2 | 16 | SE(IR) ≈ 1/√years |
| Chance one of 20 dead strategies backtests at t ≥ 2 | 64% | why the rules were hashed first |
Grinold & Kahn, Active Portfolio Management (2000), ch. 5–6, 12, 17. The 64% figure is theirs: given twenty informationless strategies, the probability that at least one shows a t-statistic of 2 is 64%, which is why a backtest that survived a search is worth so much less than it looks.
Risk systems are judged on the trades they stop, and almost nobody prices them. The reason is structural: transaction-cost analysis is built from executed orders, so it can only ever see the trades that happened. Grinold & Kahn call the rest censored data — "the record shows trades, not orders placed, and certainly not orders not placed because the cost would be too high" — and Wagner's studies of institutional desks found the opportunity cost of trades never made often dominates every cost that is measured. Harris lists missed-trade opportunity cost as one of the three components of transaction cost, and the hardest to see.
So a refusal log that records only the reason cannot answer the one question worth asking of a risk system: did the refusals help? Every veto in this system now carries the quotes, strikes and credit it refused, and each one is settled afterwards against where the underlying actually closed. For a defined-risk spread held to expiry that needs no model at all — the result is fixed by one number:
P&L per share = credit − max(0, Kshort − ST) + max(0, Klong − ST)
Refusals that would have lost money are money the gate saved. Refusals that would have made money are the gate's cost, and we report those with the same prominence. The arithmetic is verified against hand calculations in tools/test_opportunity_cost.py; the ledger is results/opportunity_cost.json. With a handful of refusals it is descriptive, not inferential, and it is labelled that way.
Standard practice in equity microstructure is that the microprice — the mid weighted by the size resting on each side — (Vbid·Pask + Vask·Pbid) / (Vbid+Vask) — is a better estimate of the true price than the plain mid, because it leans toward the side that is thinner and therefore likelier to be consumed next. Nobody appears to have checked it on retail options quotes, because checking it requires the true NBBO and the free feed does not publish one. We had 64 contracts where we had bought the truth.
| Sample | n | RMSE, plain mid | RMSE, microprice | paired p |
|---|---|---|---|---|
| All probed contracts | 64 | $0.1406 | $0.2289 | 0.30 |
| Excluding the single widest quote | 63 | $0.1417 | $0.1439 | 0.73 |
| Contracts our gate would trade (≤$0.20 wide) | 48 | $0.0790 | $0.0816 | 0.15 |
The microprice is not better, and on the contracts we would actually trade the two are indistinguishable. There is a mechanism for this rather than a shrug: the microprice can only ever place its estimate inside the quoted bid and ask, and we had already measured that the free feed's width is wrong — too narrow in liquid contracts, too wide in illiquid ones. An estimator defined inside a mis-stated interval inherits the mis-statement. Displayed-size imbalance may well carry information in options; this says only that reading it through a mis-stated spread does not recover the true mid. We keep the simpler estimator and publish the null: results/microprice_study.json.
Every submission in this contest reports a paper profit and loss, and each of those numbers inherits whatever the paper venue assumes about liquidity. Nobody appears to have measured what that assumption is, so we did.
Theory is specific about what a real market does. Impact grows as the square root of size traded — Grinold and Kahn derive it from the liquidity supplier's inventory risk, Loeb's 1983 block-bid data fits the curve, and BARRA's own fitting put the exponent at one half. Displayed depth at the touch is small, so a large marketable order walks the book by construction. We sent one contract, then 49, then 196 — the last of those 1.34× the entire displayed offer — and recorded what came back.
| Size | Displayed at the offer | Multiple | Quoted ask | Filled at | Round trip |
|---|---|---|---|---|---|
| 1 | 67 | 0.01× | $1.24 | $1.21 | $0.00 |
| 49 | 178 | 0.28× | $1.19 | $1.26 | +$0.02 |
| 196 | 146 | 1.34× | $1.30 | $1.27 | +$0.01 |
Two things here survive any objection. The round trip is a matched pair — buy and sell the same contract seconds apart — so market drift cancels out of it. In a real book a round trip costs at least the quoted spread, because you buy at the offer and sell at the bid; that is as close to a law as microstructure has. Here not one round trip cost anything, and two of them paid us. And the largest order took 1.34 times the entire visible offer and filled three cents inside it, where a real book would have exhausted the offer and walked up to worse prices.
The square-root law does not operate in this venue. Order size is free here in a way it never is in a market. That is not a complaint about Alpaca — a paper engine is a teaching tool, not a market simulator — but it does mean every paper profit in this contest, ours included, was earned somewhere that gives away liquidity a real market sells.
What it does not establish: a size-impact curve. Three rungs, and on 0DTE contracts the quote moved several cents between our snapshot and the fill, so the per-rung slippage figures mix size with ordinary market movement. We do not read them as impact. The round trip and the depth multiple are the two measurements that survive that confound, and they are the only two we draw a conclusion from.
Raw rungs in results/size_ladder.jsonl, analysis in tools/size_ladder_analyze.py. The probe cost nothing and its profit is excluded from every trading figure we report — see the note on attribution below.
One paper account has carried two different activities: the agent trading its strategy, and research probes that deliberately buy and sell to measure how the venue behaves. The broker adds them together. We do not.
A probe that happens to end up ahead is not a trading result by any reading, and blending it into one number is the exact thing our reporting rules forbid. Every fill is attributed to the thing that caused it, using the agent's own journal as the record of what it traded; anything else on the account is research and is reported separately, whichever way it lands.
Run it yourself: python3 tools/pnl_attribution.py.
There is a mature literature of optimal execution, and we used almost none of it. That is a choice, not an oversight, and the reasons are specific.
The pattern is deliberate: where a technique needs a parameter we cannot measure at this budget, we left it out and said so, rather than shipping an unvalidated estimate with a confident name.