Historical testing
Does MakeItBy make better choices than simply taking the earliest flight?
We replayed past flight choices using only information that would have been available before each trip. MakeItBy chose a flight, the comparison strategy chose the earliest eligible flight, and we checked which one actually arrived before the deadline.
Net additional deadline successes
4,957 more historical choices made the deadline, net, than always taking the earliest eligible flight.
This subtracts the cases where the earliest flight succeeded and the MakeItBy choice did not.
The simplest result
When MakeItBy's option had a clearly better historical record, its choices missed the deadline less often.
The chart compares missed deadlines across the same historical decisions. Lower is better.
What the actual arrivals looked like
MakeItBy often chose a later scheduled arrival, but fewer of those choices ended up past the deadline.
That shifts the green curve closer to the deadline. The key comparison is how much of each curve falls on the missed-deadline side.
Head-to-head outcomes
When only one of the two choices made the deadline, MakeItBy won more often.
Each group shows how much more often the MakeItBy option historically made the deadline than the earliest flight. As that historical difference grew, the gap in actual outcomes also grew.
Why the app is selective
Small historical differences are not enough to justify giving up an earlier arrival.
An earlier version could choose a later flight even when its historical record was only slightly better. Live testing exposed why that was not enough, so the current app requires a clearer difference before recommending a later option.
Consistency across time and deadlines
MakeItBy performed better in every year-and-deadline test.
We tested 3 separate one-year periods and 3 arrival deadlines in each period. The current recommendation rule performed better than the earliest-flight strategy in all 9 combinations.
A separate confirmation test
The original recommendation rule also worked on a year that was not used to develop it.
We developed the original recommendation rule using later historical periods, then applied it without changes to July 2023 through June 2024. It beat the earliest-flight strategy at all 3 tested deadlines.
The current one-extra-success-per-100 requirement was added later
That stricter requirement was added after live testing exposed a weak recommendation. Historical results support the stricter version, including better results in all nine year-and-deadline groups, but that exact requirement has not yet been tested on a future period that was unavailable when it was chosen.
How the historical test worked
The test compared decisions, not just model scores.
Use only earlier records
For each past travel date, the model used only flight history from before that date.
Make two choices
MakeItBy selected a flight. The comparison strategy selected the earliest flight that could meet the deadline.
Check the real outcomes
We checked whether each selected flight actually arrived before the deadline.
How to read the result
Historical reliability helps compare choices. It does not forecast the exact flight you will take.
If MakeItBy shows that one option has a stronger historical deadline record, that is evidence for comparing the scheduled choices available to you. It is not a guarantee that the selected flight will make the deadline.
The earliest flight is already a strong baseline
MakeItBy is not trying to replace an obviously bad default. Its value is finding the cases where historical reliability supports a different choice.
No recommendation is a valid result
If the flights are too similar or there is not enough comparable history, the app should not force a choice.
Day-of-travel conditions still matter
Weather, air-traffic restrictions, maintenance, cancellations, gate changes, and other current conditions can change the risk after this comparison is made.