In the language of trading, the word signal carries a quiet exaggeration. It implies something received rather than constructed, a message already present in the data, waiting to be read by anyone patient enough to look. The working reality is closer to the opposite. A signal is the end of a long process of attrition: the small part of an idea that survives after cleaning, cost, slippage, crowding, and regime change have each removed their share. What finally reaches a live trading limit is not a discovery. It is a survivor.
This distinction is not pedantry. It changes what a research desk should spend its time on. If signals are discovered, the work is search: try more data, more features, more models, and edge accumulates. If signals are survivors, the work is mostly attrition management: deciding what to subtract, in what order, and how much certainty to demand before committing capital. The first view rewards cleverness. The second rewards a protocol. This research argues for the second.
Section 01The observation
Every candidate begins as a noticed regularity. A spread tends to widen at a particular hour; a security drifts ahead of a scheduled release; one instrument seems to lead another by a few seconds. The notice is real, but the data underneath it is rarely as clean as the chart implies. Raw market data carries survivorship gaps, unadjusted corporate actions, stale or out-of-order timestamps, venue-specific quirks, and the silent assumption that a printed price was actually executable. A regularity that looks like structure is often an artefact of how the data was assembled.
So the first discipline is to distrust the observation in proportion to how much you like it. The more an apparent pattern flatters a thesis you already hold, the more aggressively it should be checked against reconstruction error, look-ahead leakage, and the possibility that the effect lives entirely in the half-second between a quote and a fill. Most promising patterns do not survive this stage, and that is the stage working correctly.
Section 02From notice to hypothesis
A notice becomes a hypothesis when it is stated in a form that can be wrong. Not "momentum works," but "the sign of the past k-day return predicts the next day's return in this universe, after costs, with a stability that does not depend on a single regime." Feature construction is where the claim acquires a precise shape. An economic mechanism is not proof, and genuine effects are sometimes discovered before they are fully explained, but mechanism provides constraints that help distinguish a durable relation from an unrestricted search result.
Here the research process meets one of its central identification problems. When a desk tries hundreds of features and retains those that test well, conventional significance thresholds no longer have their usual interpretation. Harvey, Liu and Zhu, surveying the catalogue of published return factors, argue that heavy mining of the cross-section requires a higher hurdle than the textbook t > 2, closer in their calibration to t > 3. Bailey, Borwein, López de Prado and Zhu make the operational point: as the number of configurations grows, an apparently strong backtested Sharpe ratio becomes increasingly likely under selection alone. Unless the search count is recorded, a reader cannot separate skill from specification search. The backtest is therefore evidence only in conjunction with the process that produced it.
A mechanism does not validate a feature. It narrows what the feature is allowed to mean, which makes failure more informative and extrapolation less arbitrary.
Research log, January
Section 03The attrition
Suppose a feature clears the significance hurdle and carries a plausible mechanism. It now faces the frictions that turn paper edge into realised edge, or into nothing. The first is cost. A strategy measured at the midpoint quietly assumes it can trade at a price no counterparty will offer. Crossing the spread, paying impact, and bleeding through slippage are not rounding errors; for short-horizon signals they are frequently the whole position. The optimal-execution literature, beginning with Almgren and Chriss, frames this as an explicit trade-off: trade quickly and pay market impact, or trade slowly and accept the risk that the price moves away while you wait. Either way, the cost is structural, and it scales with the very thing a good signal wants to do: size up when the opportunity is largest.
The second friction is decay. A signal's expected return is not a constant of nature; it is a function of how many people are trading it. The moment an edge becomes known, capital arrives, the inefficiency compresses, and the half-life shortens. The third is regime change: the relationship that held across one volatility environment, one rate regime, one market structure, can simply stop holding when the conditions that generated it change. None of these are model errors. They are properties of trading in a system that adapts to being traded.
Method · The research ledger
Selection bias cannot be repaired by subtracting one universal penalty. The minimum discipline is to preserve the search path so that the reported statistic can be interpreted in context.
for each candidate:
record(data_vintage, universe, feature, parameters, benchmark)
record(all_trials, including rejected configurations)
estimate(out_of_sample_return, turnover, cost, capacity)
report(sensitivity_across_periods_and_specifications)
Deflated performance statistics and false-discovery controls become meaningful only when the number and dependence of trials are visible.
Section 04The protocol is the product
If apparent edge is abundant and durable edge is rare, the scarce asset on a research desk is not ideas but a process that reliably distinguishes the two. That process has a recognisable shape. It keeps a research log that records every configuration tried, so the denominator of the search is never hidden from the person reading the result. It reserves a genuinely untouched out-of-sample period and refuses to spend it casually. It paper-trades before it risk-trades, so that the gap between modelled and realised fills is measured rather than assumed. It sizes risk to the uncertainty of the estimate, not to the seduction of the backtest. And it monitors live performance for decay, treating a fading signal as expected news rather than a personal failure.
The backtest is not the evidence. The process that produced it is the evidence.
Figure 1 is therefore an evidentiary lifecycle rather than a production line. Reconstruction establishes that the observation existed. Multiple-testing control establishes how surprising it is relative to the search. Execution modelling establishes whether it can be implemented. Live monitoring tests whether the proposed mechanism remains compatible with new data. The final research object is not the feature alone but the feature, its provenance, its feasible trading policy, and the conditions under which belief in it should be reduced.
Caveats
- This is a description of a research process, not a recipe for a profitable strategy. No feature, threshold, or sizing rule here should be read as a recommendation.
- The significance hurdles cited are estimates from the academic literature on a specific dataset of equity factors; they are useful as intuition, not as universal constants.
- Real desks face constraints this research abstracts away (capacity, financing, borrow availability, and operational risk), any of which can dominate the statistical story.
This research is analysis and commentary for general information. It is not investment advice, an offer, or a solicitation, and it contains no price forecasts. Where the text reports findings from cited sources, it distinguishes those facts from the author's interpretation.
References & notes
- Harvey, C. R., Liu, Y., & Zhu, H. (2016). “… and the Cross-Section of Expected Returns.” The Review of Financial Studies, 29(1), 5–68. Source for the multiple-testing critique and higher evidentiary hurdle in the published factor literature.
- Bailey, D. H., Borwein, J. M., López de Prado, M., & Zhu, Q. J. (2014). “Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance.” Notices of the American Mathematical Society, 61(5), 458–471. Consulted for the operational problem of selecting the best result from an undisclosed search.
- Almgren, R., & Chriss, N. (2000). “Optimal Execution of Portfolio Transactions.” Journal of Risk, 3(2), 5–39. Source for the expected-cost versus timing-risk framework.