OpenAI wants a four-day week funded by AI productivity gains. The most-cited evidence for it is a trial that mostly happened before ChatGPT existed.
Written for the owner who just saw OpenAI's proposal and the UK's four-day-week numbers quoted in the same paragraph. The numbers are real. They aren't evidence of what's being claimed, and the arithmetic that would actually prove it hasn't been run.
OpenAI's April 6 policy blueprint, "Industrial Policy for the Intelligence Age," proposes a long list of ideas for managing AI's economic disruption — automation taxes, a public AI wealth fund, affordable access for small businesses — and one that travel-sized coverage keeps isolating: an "efficiency dividend," incentivizing employers and unions to run time-bound 32-hour, four-day pilots at full pay, with the saved hours banked as permanent shorter weeks or paid time off once routine workloads actually shrink. Coverage of the proposal treats the idea as close to proven, and almost always reaches for the same supporting evidence to say so.
The trial everyone reaches for
That evidence is the UK's four-day week pilot: 61 companies and roughly 2,900 employees moved to a 32-hour week on full pay from June to December 2022, coordinated by 4 Day Week Global and the Autonomy Institute with researchers from the University of Cambridge and Boston College. The results, published in February 2023, were genuinely strong: revenue stayed roughly flat during the trial and was up about 35% on the same period the year before, staff turnover fell 57%, and 56 of the 61 companies kept the four-day week afterward, with 18 making it permanent. It's a real result, and it gets cited constantly as proof that AI can fund a shorter week without anyone losing output.
The gap in that citation
The trial ran from June to December 2022. ChatGPT launched to the public on November 30, 2022 — five of the trial's six months happened before it existed at all, and the sixth was the month it launched in. None of the 61 companies were running a generative AI assistant in daily work; there wasn't one to run yet. The productivity-neutral four-day week in that trial came from the lever every four-day-week study before and since has found: fewer meetings, protected deep-work blocks, and low-value work trimmed by the people doing it, not software. That's a real and useful result about reorganizing a work week. It is not evidence that AI-driven time savings can fund one, because no AI was present to generate any.
What the actual AI productivity numbers measure
The gains that are genuinely tied to AI exist, and they arrived later, but they describe something narrower than a shortened role. Brynjolfsson, Li and Raymond's NBER study of 5,172 customer-support agents found a generative AI assistant raised issues resolved per hour by 15% on average — one task, in one job, with the biggest gains going to the least experienced agents and almost none to the most experienced. A Boston Consulting Group study of 758 consultants using GPT-4 found they finished 12.2% more tasks 25.1% faster on work inside what the researchers called AI's "jagged frontier" — and did worse on tasks outside it. Stanford's 2026 AI Index puts the range at 14% to 26% across customer support and software development specifically. Every one of these numbers measures one task against its own baseline, with real limits at the edges. None of them says a 20% gain on a narrow task frees a fifth of anyone's whole working week — that's a different, untested claim riding on the first one's credibility.
The arithmetic to run before cutting a day, not after
ROI here is arithmetic, not a multiplier borrowed from someone else's study. Before shortening anyone's week, pick the one task AI is actually supposed to be shortening on this team, and time it — before AI and after, on real instances, under the real deadlines the work actually has, not a demo run with no client waiting on the other end. That's the only way to find out whether the gain is 15% or 2% on this specific task, for this specific person, and whether it's the kind of time that adds up to a freed afternoon or just a slightly calmer Tuesday. The team that skips this step is doing what most of this year's coverage did: assuming a number measured somewhere else applies here, at full strength, before checking.
The freed hours don't stay freed on their own
Say the arithmetic holds and a task really does shrink by a fifth. The hours it frees aren't automatically banked into a shorter week — they quietly refill with whatever else is sitting in the inbox unless someone named is checking, a few weeks later and again a quarter later, that the task is still taking less time and the freed hours haven't crept back into the same five-day shape. The UK trial's own 92% continuation rate held up precisely because those companies kept measuring what had changed instead of declaring victory on day one. A four-day week funded by an unverified AI estimate, with no one checking whether the estimate is still true in March, is a policy on paper and a five-day week in practice within two quarters.
None of this is an argument against a shorter week — a genuinely freed afternoon is worth having. It's an argument against funding the decision with someone else's 2022 trial and someone else's task-level study, instead of the one measurement that would actually tell a specific team whether it's earned the hours it's about to spend.
This is a position, not a critique of the underlying studies, which are credible within their own scope. The timeline point about the UK trial and ChatGPT's launch date is independently verifiable and, to my knowledge, isn't raised in coverage pairing the two.
If you're about to pilot shorter hours on an AI estimate, the methodology page walks through how to run the before/after measurement on one real workflow before anyone touches the calendar.