How Expedia predicts what a customer is worth over the next year
Not every customer is worth the same to chase or keep. Expedia built a model that predicts a year of future spend per person, across eight brands, and it is a practical template for any business with a CRM.

Some customers book once and vanish. Others come back every few months for years. If you spend the same effort acquiring and keeping both, you are wasting money on one and under-investing in the other. Customer lifetime value (CLV) is the attempt to tell them apart: how much is a given customer likely to be worth in the future?
In 2023 Expedia's data science team described how they predict it at scale in Expedia Group's Customer Lifetime Value Prediction Model. They define CLV as a customer's future cash flow over a long horizon, a year and beyond, and use it to steer acquisition, retention, incentives, marketing, and where to invest. The model runs across eight brands (Expedia, Hotels.com, Vrbo, Orbitz, and others) and five lines of business, is retrained monthly, and refreshes predictions daily for hundreds of millions of customers.
Why they skipped the textbook formulas
There are classic ways to estimate lifetime value: cohort analysis, RFM (recency, frequency, monetary) scoring, and probabilistic "buy-till-you-die" models. Expedia rejected all of them for this job, for concrete reasons: they require you to predefine segments, they ignore signals that are not purchases (like email engagement), and several of them need a customer to have a repeat-purchase history, which means they simply fail for new customers, exactly the ones you most want to understand.
So they treated it as a supervised machine-learning problem instead: learn the relationship between what we know about a customer today and what they actually spent afterward.
How they built it
- The target. For each customer they pick a cutoff date, use only what was known before it as input, and predict the next 12 months of gross cash flow after it. Separate multipliers scale that down to net value, accounting for cancellations.
- The model. CatBoost, a gradient-boosted decision-tree method. They chose it for handling high-variety categorical data, tolerating missing values, training fast, and capturing nonlinear interactions.
- The features. More than 200 inputs in two families: Bookings (recent purchase detail plus historical aggregates) and Engagement (email-marketing clicks, loyalty tier). The engagement signals are the part the textbook formulas throw away.
- The segments. Rather than one model for everyone, they split customers into 5 geographic regions crossed with recency and frequency groups, giving 30 segments, each with its own CatBoost model, and clip extreme targets at the 99.9th percentile so a few outliers do not dominate.
Measuring a number you cannot show
Here is the interesting part. Expedia would not publish the actual accuracy numbers (the chart axes are hidden for confidentiality). Yet they still validate the model rigorously, and the way they do it is the lesson.

They separate two questions. First, ranking: does the model correctly sort customers from least to most valuable? They measure that with a Lorenz curve and a Gini score, and show it beats a simple "assume next year looks like last year" baseline. Second, calibration: are the predicted values close to reality on average? They check that with calibration plots, where points should sit near a 45-degree line. You can prove a model is useful on both counts without ever publishing a single dollar figure.
They were also honest about what was hard: the pipeline was slow until they optimized it (dropping low-impact features, replacing expensive sorts with simple date math) for roughly 10x speedups, and COVID-era travel behavior was so irregular it hurt accuracy and made clean evaluation difficult.
What has changed since
- Expedia has not published a v2 of this specific model. Its public AI work since has been elsewhere: a generative-AI trip planner called Romie in 2024 (not stated to be wired to the CLV model), and a heavy emphasis on first-party data, its own record of traveler behavior, as third-party cookies faded. Google, for its part, abandoned its plan to remove third-party cookies from Chrome across 2024 and 2025, which only reinforced the value of first-party data.
- The wider field has moved on two fronts: deep temporal models for lifetime value (a 2025 system deployed at Douyin reported cutting error by a few percent), and, in 2026, early "agentic" systems that use language models to automatically build and repair the modeling pipeline itself. The hand-built gradient-boosted approach is still a strong, sane default, but the tooling around it is changing.
Why it matters for your business
You do not need Expedia's scale to use this. Any business with repeat customers and a CRM can:
- Predict forward value, not just past value. RFM tells you who was good; a model trained on outcomes tells you who will be good, including brand-new customers.
- Use engagement, not only purchases. Opens, clicks, chat activity, and loyalty status carry real signal, especially before the first or second purchase.
- Segment where behavior genuinely differs, then model within each segment, instead of forcing one model to explain everyone.
- Validate by ranking and calibration. Even if you never share the numbers, you can prove the model sorts customers correctly and predicts realistic amounts.
The payoff is concrete: spend acquisition budget on the customers likely to be worth it, aim retention offers at the valuable ones about to lapse, and stop discounting people who would have stayed anyway.
How we would build it today
For an Indonesian retailer, hotel group, or dealership, we would start with a gradient-boosted model predicting each customer's next 6 to 12 months of spend, trained on your CRM history (bookings or purchases) plus engagement signals (WhatsApp activity, email clicks, loyalty tier). Segment by the lines of business that truly behave differently. Report a Gini score for ranking and a calibration plot for accuracy, so the sales and marketing teams can trust it. Use a language model to turn messy interaction logs into clean features, and keep an honest backtest. It is the same shape Expedia used, sized to your data.