How Stitch Fix lets a machine and a human stylist pick your clothes together
Stitch Fix refused to choose between algorithms and human stylists. It built a system where the machine does the searching and the person does the judging, then fed the stylist's corrections back into the model.

Most shopping starts with browsing. You scroll, you filter, you pick. Stitch Fix built a clothing business that removes that step entirely: a client fills in a style profile, and a box of five items chosen specifically for them arrives at the door. Nobody browsed. Every item was selected by the company. That makes recommendation quality not a feature of the product, it is the product, and it forced Stitch Fix to answer a question most retailers can avoid: when a machine and a human disagree about what to send, who wins?
Their answer, laid out across the public Stitch Fix Algorithms Tour and years of engineering write-ups, is that the question is wrong. Neither wins, because they are doing different jobs.
The problem: neither the model nor the stylist is enough
Stitch Fix stated the trade-off plainly in an early post. A machine is superb at structured data: it can score fit, color and price against billions of past client-item interactions and rank tens of thousands of items in a blink. It is weak at the unstructured, human part: reading a Pinterest board, understanding a note like "I need outfits for a cruise next month," or making five items feel like a coherent, flattering set.
A human stylist is the mirror image. Excellent at taste, context and empathy. Hopeless at mentally searching a warehouse of inventory for the single best fit-probability across millions of clients. Their conclusion: "Each piece contributes distinct value to the overall selection process and exclusive focus on either one would be incomplete." So the system is built as a relay, not a contest. The machine narrows, the human chooses.
How they built it
A client is a learned position, not a fixed segment. Instead of dropping people into buckets like "classic" or "trendy," Stitch Fix learns a coordinate for every client and every item from the response matrix of who kept or returned what. Early on this was matrix factorization; the axes it discovered turned out to be human-readable style spectrums that stylists and buyers could actually use.

That model is fed constantly by Style Shuffle, a thumbs-up or thumbs-down game inside the app. By 2022 Stitch Fix reported it had gathered around 10 billion ratings, roughly 4.5 million a day, letting the system update its read of your taste within a day or two of you tapping.
One model instead of a sprawl of them. By 2022 Stitch Fix had accumulated a separate model per business line and region: Fix, Freestyle, Men's, UK, each its own thing to maintain. They replaced the pile with a single unified client embedding, described in their Client Time Series Model post. One learned representation of each client, built from their ordered history of checkouts, profile edits and Style Shuffle taps, feeds many predictions at once. Data from a US Fix can now improve a UK Freestyle recommendation, because it is all the same client vector underneath. That is what the diagram at the top of this page shows: many inputs, one embedding, many outputs.
Then the human takes over. The model hands the stylist a ranked shortlist, not a finished box. The stylist makes the final call and writes the personal note, the part no model does convincingly. Crucially, Stitch Fix does not let that human judgment evaporate. It captures it.

Every day, stylists score a random sample of algorithm-generated outfits against clear criteria, flagging where the machine went wrong. Those judgments become training labels for a model that predicts outfit problems before a client ever sees the box. Stitch Fix reported that pairing algorithmic recommendations with stylist-built outfits lifted their internal quality measure by 14 percent. A separate system fine-tuned a language model on a few thousand stylist-labelled examples to read clients' free-text feedback ("soft, stretchy, patterned") and quietly lower the odds of sending disliked attributes again, with stylists still able to override it.
What has changed since
- Stitch Fix's engineering blog went quiet after 2023, so there is less public detail on the newest internals. What is public is the business context: revenue fell about 16 percent in fiscal 2024 and about 5 percent in fiscal 2025, then grew 6.4 percent in fiscal 2026, its first annual growth after several down years. Read that as the environment the ML now runs in, not as proof any single model worked; the company does not draw that causal line, and neither should we.
- The human side got smaller. Stitch Fix reported roughly 5,100 stylists out of about 8,000 employees in fiscal 2019, and about 1,710 stylists in fiscal 2025. The human-in-the-loop design held, but with far fewer people in the loop.
- In late 2025 the company launched Stitch Fix Vision, a generative-AI tool that renders images of a client in recommended outfits from their own photos. It is the same pattern extended: the machine generates, and it is offered as an aid, not a replacement for the stylist's choice.
Why it matters for your business
You do not sell clothing in boxes, but the core idea is portable to almost any business that recommends or matches things to people:
- Split the work by strength, not by ideology. Let software do the brute-force part, ranking and searching a large catalog fast, and let your people do the judgment, taste and relationship part. Arguing about "AI versus humans" usually means you have not decided which job each is for.
- Represent each customer as a learned position, not a fixed label. "Segment 3" is a guess. A vector learned from real behavior places even a brand-new customer sensibly and drifts as they change.
- Capture your experts' corrections. The most valuable and most wasted asset in most companies is the moment a good employee overrides the system. Log that override as a label and your model gets smarter every day instead of repeating the same mistake.
How we would build it today
For an Indonesian retailer, salon, clinic or property agency, we would start small and human-first. Build one model that ranks your catalog or listings for a given customer from your own history of purchases, bookings and returns, and surface a shortlist to your staff rather than auto-sending anything. Give the staff a one-tap way to accept, swap or reject each suggestion, and store every one of those actions as training data, because that is your stylist-in-the-loop. Turn each customer's history into a simple learned embedding rather than a rigid segment. Use a language model to read messy free-text signals, WhatsApp chats, review notes, form comments, into structured preferences, with a person able to override it. Then backtest honestly: train on the past, and check the recommendations against what customers actually chose next. It is the same shape as Stitch Fix, sized to your data and your team.