Every tested trading idea touches two kinds of data. One is the data used to build and adjust the rule. The other is data the rule has never seen. The first is called in-sample and the second out-of-sample. Most of what goes wrong in testing comes from mixing them up.

What in-sample means

In-sample data is whatever you looked at while making choices. If you picked a moving-average length because it gave the best result over the last three years, those three years are in-sample. The result on that stretch describes how well the rule was fitted. It says little about how the rule behaves on anything else.

This is the mechanism behind overfitting. Give a rule enough adjustable parts and it can be shaped to match any stretch of history, noise included. The in-sample result will look excellent for exactly that reason.

What out-of-sample means

Out-of-sample data is held back. You build the rule on one stretch, lock every setting, then run it once on the stretch you kept aside. Because no choice was made with that data in view, the result is a fairer picture of the rule meeting something new.

The usual finding is that the out-of-sample result is worse. That is expected. Some of the in-sample result was the rule fitting noise, and noise does not repeat. A small drop suggests the rule captured something that persisted across both periods. A collapse suggests most of it was fit.

How the hold-out gets spoiled

The hold-out works once. Run the rule on it, dislike the result, change a setting and run it again, and the held-back data has become part of the tuning. Do that ten times and it is in-sample in everything but name. Nothing in the software stops this. Only the person running the test can.

There are quieter leaks. Choosing which market to test after seeing which one trended is a choice made with the full data in view. So is choosing the start date. So is reading about a strategy that is popular because it did well recently, then testing it over the same recent period. Each one lets information from the test period into the design, which makes it a cousin of look-ahead bias.

Splitting the data

The simplest split is by time: an early block to build on and a later block to check on. Time order matters because markets change, and a rule will only ever be used on data that comes after it was built. Walk-forward optimization repeats that split many times, rolling the build window and the check window forward together.

The last out-of-sample test is live. Running a rule forward on paper or with small size is the one test that cannot be contaminated, since the data did not exist when the rule was written. Backtest versus forward test covers that step.

What to ask of any result

When someone shows a tested result, three questions sort it. Which data was used to choose the settings. Which data was never touched until the end. How many times the end was visited. A result with no answer to those is an in-sample result, whatever it is called.

None of this proves a rule will hold. An honest out-of-sample test only removes one common way of fooling yourself. The market is free to change after the test ends, which is the subject of regime change and old statistics.