UX Case study
Adding pay later to a hotel booking flow
Findhotel already had rooms in its portfolio that a traveller pays for at the property rather than at checkout. What it did not have was a way for anyone to choose that whilst booking. This is the first version of that choice, and the five-week experiment that measured what it was worth.
The choice, where it ended up: on the room selection page, before checkout. Not where I wanted it, and the reason why is further down.
Problem
Why this project?
Findhotel is a smaller player in accommodation, working the localised markets that Booking, Expedia and TripAdvisor tend to leave alone, and much of its portfolio is distress inventory: rooms a hotel expects not to sell and cannot list itself without being deprioritised by the larger platforms. Paying at the property rather than at checkout is one of the things that can distinguish a platform like that from the ones above it.
So the business case was never in question. What was in question was everything else. Letting a traveller pay later meant acquiring the right hotel portfolio, restructuring how bookings were handled, and then, at the end of all of it, giving somebody a way to choose it. The last part is the one this case study is about, and it was harder than it sounds.
The real problem was that pay later could go a thousand directions. Everyone had a view on how it should look and on what belonged in a first version. Introducing a feature across several scrum teams is difficult in any organisation; it is considerably harder when the feature is expected to move the conversion rate of the company's main product, because then everybody is right to have a view. Something had to settle it.
One thing this page cannot tell you is what a traveller said they wanted, because nobody wrote that down. The case that exists on paper is the commercial one, and it would be dishonest to reconstruct the other half now.
Goals
Desired outcome
We knew the feature would be welcome. What nobody knew was by how much, and that gap is what the first version was designed around.
Those three are one goal wearing three hats. A minimum viable product here was not a way of doing less work; it was the only way to get a number, and a number was the only thing that would buy the argument for a second round. Everything I would have preferred to design was waiting on it.
Team & audience
Who built this, and who for?
Team's setup
Our product owner brought the business case and the priorities. Our data scientist brought the numbers the decisions were made against, and designed the A/B test itself, which matters more than it sounds: a test designed by the people who want a particular answer is not a test.
A dedicated copywriter wrote the words. An external agency translated them. I was the only designer on the feature, and the developers were a scrum team inside an organisation that had several.
My responsibilities
Benchmarking the competition, running the team's ideation and scoping sessions, wireframing, building a prototype low-fidelity enough to argue with, guerrilla testing it, turning the result into a development plan, and drawing the high-fidelity screens against the design system.
After that: keeping the copy pointed at what the test was measuring, testing the build by hand against every scenario we had agreed on, watching real sessions once it was live, and writing the whole thing up for colleagues who were never in the room.
Audience
For whom are we building this?
English speakers first, and that was a decision made for us. The translation system at the time was slow enough that pushing new copy through it in the weeks we had was not realistic, and the copy was expected to go through several rounds of refinement anyway.
So the first weeks of the test shipped in English only. An English-speaking traveller saw the whole flow in their own language. A Korean, Thai or Danish traveller saw the rest of the site in theirs and the pay-later terminology in English, which is a genuinely odd thing to hand somebody in the middle of a purchase. What made it defensible was the shape of the audience: American users are by a wide margin the largest group on the platform, so the people who would see the feature whole were also the people most likely to see it at all.
Scope and constraints
Limitations
Three constraints shaped this more than any design decision I made.
The first is the one above: mixed-language copy in the middle of a checkout, for everyone outside English.
The second was the search page. If a hotel supported pay later, the search results ought to say so, and that is where a traveller decides which hotel to open. The search page belonged to another part of the organisation and could not be changed in time for the test. The result is worth stating plainly, because it comes back at the end of this page: a traveller searching on Findhotel would see pay-later offers from Booking, Agoda and Expedia in the results, and none from Findhotel, even where Findhotel had them.
The third was the checkout page, where a technical limitation in the architecture ruled out anything but minimal changes. That took several designs I preferred off the table and left exactly one place where a traveller could make the choice: the room selection page. Asking somebody to commit to a payment schedule before they have reached the checkout is odd at best. It was the version that could ship, and shipping something measurable was the whole point.
Process
Step by step description
What follows is one version of one feature, not the history of pay later at Findhotel. It is the first one, and the one that had to earn the rest.
Scenario mapping
Before anything else, a rough map of what a traveller could do and what would happen when they did. Not a flow of the happy path: a spread of the outcomes, including the ones nobody wants.
The reason to draw it first is that the development plan later is only as good as this. Mapping the scenarios up front is how edge cases get handled as design decisions rather than as bugs found in the last sprint. It does not catch everything, and this one did not.
Every interaction, and what follows from it. The edge cases are the point of this drawing, not the middle.
Competitor benchmarking
Every major player in accommodation, and several minor ones, taken apart screen by screen. The point was a set of findings the team could act on: what the industry had settled into, where the gaps were, and which of our arguments had already been answered by somebody with more traffic than us.
It had a second effect I did not plan for. Working through other people's flows turned up more questions about the industry than it answered, and several of those went out to other members of the team to research. That is what looking at how others do it is for: it is exploratory, and the useful part is often the question rather than the copy of the answer.
Booking.com and Expedia. The two largest platforms in accommodation, and the two worth studying hardest.
Ideation and scoping
I ran the ideation sessions, and the benchmark was the reason they worked. A room full of people with strong opinions about a feature is a difficult room; a room full of people who have just been shown what five competitors actually shipped is a different one. The material did the arguing.
Scoping ran alongside, with the product owner and the data scientist. Two questions, kept separate on purpose: what would the ideal version of this be, and what is the smallest version worth putting in front of anybody. Holding both meant the cuts were recorded as cuts rather than lost, which is what makes a second round possible.
Wireframes, and testing them in the corridor
The wireframe became a clickable prototype, and its job was to put two competing concepts into a form somebody could hold. An opinion about a diagram is cheap. An opinion about a thing you have just clicked through is worth having, and it arrives earlier.
Deliberately low fidelity, because the audience was stakeholders and anyone in the organisation with an interest in the outcome, and a polished screen invites a conversation about the polish. Then guerrilla testing, to knock out the obvious problems before they cost anybody a sprint, and to surface the questions the concept raised. Several of those went back to the team to research properly.
The screen the whole feature turns on, in the form it was first argued about. Two ways to pay for the same room at the same price, one under the other, and nothing else on the screen to argue about.
The development plan
One drawing holding the happy path, the main edge cases and the screen each of them lands on. From that the product owner and the developers could write tickets without coming back to ask what happens when.
It also settled something the design could not decide on its own. Our architecture ran on Expedia's data for this version, which meant conforming to Expedia's rules for how a booking is processed after checkout. So I read those rules and redrew them as a decision flow a non-technical colleague could follow, which turned out to be as useful to the developers as to anyone else: it made visible which sub-features we were inheriting and which we were choosing not to implement yet.
Every screen that had to change, and what had to change on it.
What happens to a booking after checkout, on the rules we were inheriting. Cropped here; it opens in full.
High fidelity screens
Once the tickets existed, the screens were drawn against Findhotel's style guide and component library, so that a front-end developer was implementing a layout rather than interpreting one. Paddings, colours, margins and icons all came from the system. The key screens for desktop and mobile were published in Zeplin for the developers to work from.
Testing it by hand, before release
Once developers had something unreleased, I worked through the scenarios we had agreed for that sprint by hand. This is where the scenario map and the scoping diagrams earn their cost: without them, "did we test everything" is a matter of memory.
Alongside the scenarios, the team checked the same build against every one of these, wherever they applied. It is a long list for a feature this small, and that is the honest shape of shipping a payment choice into a booking funnel that already exists.
- Device mobile, tablet, desktop
- Browser Chrome, Safari, Firefox, Edge
- Stay length a single night against several
- Room configuration one room against several
- Signed in against signed out
- Private deals present, absent, and mixed
- Availability available, sold out, price mismatch
- Tax display excluded, as in the US, against included, as in the Netherlands
- Payment breakdown prices carrying pay-at-property taxes and fees against pay-now-only prices
- Currency converted from the provider against unconverted
- Language English against German, which is longer and breaks layouts first
Watching it in the wild
Once it was live, the question stops being whether it works and becomes what people actually do with it. For that I used FullStory, which records real sessions and lets you segment by behaviour or by where somebody is. When I found a problem, a bug, or an idea worth keeping, I wrote it down. Bugs slip through, and interfaces do strange things in particular browsers; this is how you find out which ones.
Five sessions, kept because each of them shows something the numbers cannot. Two are on the A side and three on the B side of the test.
Taiwan, desktop, A side
India, desktop, B side
Ohio, phablet, B side
Qatar, mobile, A side
Spain, tablet, B side
Writing it up for everyone else
The last step, and the one that is easiest to skip. An experiment that only its own team understands has told the company nothing.
The product owner and the data scientist wrote the readable account. What I made was the pair of drawings below: the two sides of the test, side by side, so that a colleague who had never sat in one of our meetings could scroll past and see in a few seconds exactly what had been changed.
Conclusion
Results
Conversion on Expedia offers, which were the only offers in the test. That figure is already net of the higher cancellation rate below, so it is what the feature was worth after the cost of the bookings it lost.
Cancellations on pay-later bookings against pay-now ones. Not a surprise, and not free: a booking somebody has not paid for is easier to walk away from. It is the price of the conversion above.
And then the part that is less comfortable to publish. A bug in the search page meant only the top three providers were shown for each hotel listing, and Findhotel's own pay-later offers were not always amongst them. So an unknown share of the travellers in this test could not find the thing being tested.
Which means the number above is a floor rather than a measurement. The feature was probably worth more than 4.2%, and the honest version of that sentence is that we do not know how much more. The search page was fixed afterwards, but the test was never run again against a corrected one, so the ceiling is still unmeasured. What the test did establish is a direction, and a direction was enough to justify the work that came after it.