Historical testing

Does MakeItBy make better choices than simply taking the earliest flight?

We replayed past flight choices using only information that would have been available before each trip. MakeItBy chose a flight, the comparison strategy chose the earliest eligible flight, and we checked which one actually arrived before the deadline.

+4,957

Net additional deadline successes

4,957 more historical choices made the deadline, net, than always taking the earliest eligible flight.

This subtracts the cases where the earliest flight succeeded and the MakeItBy choice did not.

About 1 in 139 qualifying historical decisions produced one net additional deadline success.
9 of 9 year-and-deadline tests favored MakeItBy.
This is historical testing, not a live forecast. These results do not predict the exact outcome of a future flight and do not account for current weather, air-traffic restrictions, maintenance, or other day-of-travel conditions.

The simplest result

When MakeItBy's option had a clearly better historical record, its choices missed the deadline less often.

The chart compares missed deadlines across the same historical decisions. Lower is better.

Historical comparison of missed deadlines per 1,000 trips for the earliest-flight strategy and the MakeItBy choice.
Historical comparison only. The app only recommends arriving later when its option historically made the deadline at least one more time per 100 comparable trips.

What the actual arrivals looked like

MakeItBy often chose a later scheduled arrival, but fewer of those choices ended up past the deadline.

That shifts the green curve closer to the deadline. The key comparison is how much of each curve falls on the missed-deadline side.

Smoothed historical arrival-time distributions for the earliest-flight strategy and MakeItBy choices, centered on the traveler's deadline.
Cancellations, diversions, and flights without a recorded arrival time are not placed on the curve, but still count as missed deadlines.

Head-to-head outcomes

When only one of the two choices made the deadline, MakeItBy won more often.

Each group shows how much more often the MakeItBy option historically made the deadline than the earliest flight. As that historical difference grew, the gap in actual outcomes also grew.

Historical head-to-head outcomes showing how often MakeItBy made the deadline when the earliest flight missed, and vice versa, grouped by the size of the historical difference between the two options.

Why the app is selective

Small historical differences are not enough to justify giving up an earlier arrival.

An earlier version could choose a later flight even when its historical record was only slightly better. Live testing exposed why that was not enough, so the current app requires a clearer difference before recommending a later option.

36.8% of the original later-flight choices had less than one extra historical deadline success per 100 comparable trips. The current app declines to make the later recommendation in this range.
29.7% of the original later-flight choices had at least two extra historical deadline successes per 100 comparable trips.
about 47% of the choices retained by the current rule had at least two extra historical successes per 100 comparable trips.

Consistency across time and deadlines

MakeItBy performed better in every year-and-deadline test.

We tested 3 separate one-year periods and 3 arrival deadlines in each period. The current recommendation rule performed better than the earliest-flight strategy in all 9 combinations.

A separate confirmation test

The original recommendation rule also worked on a year that was not used to develop it.

We developed the original recommendation rule using later historical periods, then applied it without changes to July 2023 through June 2024. It beat the earliest-flight strategy at all 3 tested deadlines.

11 AM deadlineBetter than earliest
12 PM deadlineBetter than earliest
1 PM deadlineBetter than earliest

The current one-extra-success-per-100 requirement was added later

That stricter requirement was added after live testing exposed a weak recommendation. Historical results support the stricter version, including better results in all nine year-and-deadline groups, but that exact requirement has not yet been tested on a future period that was unavailable when it was chosen.

How the historical test worked

The test compared decisions, not just model scores.

1

Use only earlier records

For each past travel date, the model used only flight history from before that date.

2

Make two choices

MakeItBy selected a flight. The comparison strategy selected the earliest flight that could meet the deadline.

3

Check the real outcomes

We checked whether each selected flight actually arrived before the deadline.

3separate one-year periods
40U.S. destination airports
3deadlines in each period

How to read the result

Historical reliability helps compare choices. It does not forecast the exact flight you will take.

If MakeItBy shows that one option has a stronger historical deadline record, that is evidence for comparing the scheduled choices available to you. It is not a guarantee that the selected flight will make the deadline.

The earliest flight is already a strong baseline

MakeItBy is not trying to replace an obviously bad default. Its value is finding the cases where historical reliability supports a different choice.

No recommendation is a valid result

If the flights are too similar or there is not enough comparable history, the app should not force a choice.

Day-of-travel conditions still matter

Weather, air-traffic restrictions, maintenance, cancellations, gate changes, and other current conditions can change the risk after this comparison is made.