What a Control Group Looks Like for a Trading Rule

Halfway through the period, when a new filter appears to be helping, the only honest question is what the account would have done without it. The note orb trading review consultoriainnova publishes on this covers the baseline first, because a strategy for the opening range changes its own sample the moment a filter is added, and comparing the filtered results with the unfiltered results from last year compares two different sets of market conditions rather than two rules.
The Baseline Has to Run Concurrently

A control group in this context is a paper record of what the unfiltered rules would have produced over exactly the same sessions. Same names, same timeframe, same entries and exits, minus the change under test. Because it covers identical conditions, any difference between the two records is attributable to the rule rather than to the quarter.
It Costs a Spreadsheet, Not Money

Nothing is traded on the baseline. Each session, the entries the unfiltered rules would have taken get logged with their trigger price, stop loss and outcome at the planned exit. Ten minutes at the closing bell covers a small watchlist. That is the entire cost, and it converts an argument about whether a filter helps into a pair of numbers.
Watch the Frequency, Not Only the Average
Filters almost always improve average result per trade, because removing trades removes losers along with winners. The comparison that matters is total outcome across the period, which combines the average with how many opportunities survived. A filter that lifts expectancy while cutting the trade count in half may leave the account worse off, and only the concurrent baseline shows that clearly.
Look at What It Removed, Individually
Beyond the totals, the list of excluded trades is worth reading one by one. If the filter is removing a scattered mixture of winners and losers, it is adding noise. If it is consistently removing a recognisable kind of session, such as narrow ranges or low relative volume mornings, it is doing something real. The pattern in the exclusions is stronger evidence than the improvement in the total.
Retire the Baseline Deliberately
Running a shadow record forever is a burden and eventually it stops being read. Fix an end date when the test starts, make the decision at that date, and stop. If the filter is adopted it becomes part of the rules and gets a version number, and the next test needs a fresh baseline built against the new set rather than the old one.