SmarterArticles

Keeping the Human in the Loop

The hillside above Adavivaram has been stripped to red earth and cut into terraces, a staircase of raw laterite climbing away from the coast. Reuters journalists who walked the site in early August found work proceeding at full pace, earthmovers reshaping a slope that until recently was green. Below the cut, roughly 120 metres away according to the Human Rights Forum, sits the Mudasarlova reservoir, one of the water bodies that keeps Visakhapatnam's taps running. Within a kilometre in the other direction lies the Kambalakonda Wildlife Sanctuary, home to leopards and pangolins, and, in the assessment Reuters reported, only 860 metres from the construction site itself.

In the city below, activists and children have marched with banners reading “We cannot drink DATA”, some painting handcuffs on to the Google logo. On the Sunday before the Reuters visit, campaigners met to plan door-to-door awareness drives and beach protests. It is not the reception a 15 billion dollar investment usually gets.

This is Google's largest AI infrastructure project outside the United States: a gigawatt-scale campus, built with the Adani Group and Bharti Airtel, announced in October 2025 and intended to make Visakhapatnam a landing point for subsea cables and a node in the global machinery of artificial intelligence. Andhra Pradesh's chief minister, Nara Chandrababu Naidu, laid the foundation stone in April 2026, targeting completion by September 2028. The state calls it historic and transformational.

It is also being built in a city that does not have enough water. By the state's own accounting, cited by Reuters, Visakhapatnam receives about 410 million litres a day against a requirement of 480 million. Rationing is routine for a population of roughly 2.5 million. Mongabay India reported that Visakhapatnam district holds the lowest volume of available groundwater of any district in Andhra Pradesh, 2.12 thousand million cubic feet as of 1 April 2026. Into this basin, the state has promised the project guaranteed water for twenty years.

The question that raises is not simply whether Visakhapatnam can spare the water. It is a question about the shape of the AI economy: what it means that the physical substrate of a technology consumed overwhelmingly in wealthy countries is being poured, literally, into the ground of places that cannot afford it, and what anyone in those places is entitled to demand in return.

Why Andhra Said Yes

Start with the strongest possible version of the case for building it.

India is not a bystander in the AI economy. Mongabay India reported that the country hosts something like 20 per cent of the world's data while holding around 3 per cent of global data centre capacity. That asymmetry is not neutral. Indian data sits on foreign soil, under foreign jurisdiction, and the value generated by processing it accrues elsewhere. Every argument European governments have made for digital sovereignty applies to India with more force, because India has more to lose.

Andhra Pradesh wants to change that, pursuing roughly 6.5 gigawatts of compute capacity. India's total stood at about 1.5 gigawatts at the end of 2025, up from around 375 megawatts in 2020, with Deloitte forecasting 8 to 10 gigawatts by 2030. Visakhapatnam, with its coastline, port, subsea cable potential and engineering colleges, is the obvious place for a large slice of it.

The economics offered are considerable. State officials have put job creation at up to 188,000 across the project's ecosystem. Google has committed to new transmission lines, clean energy generation and storage, and says it is building its own renewable generation without state or central incentives. The Tribune reported the allotment of 480 acres across Visakhapatnam and Anakapalli districts. Land has been discounted, stamp duty waived, electricity and water tariffs reduced and tax reimbursed during construction, a package the News Minute totalled at around 22,000 crore rupees.

Then there is the moral argument, which deserves to be taken seriously rather than waved away. Western commentators objecting to a data centre in India are objecting to infrastructure of a kind their own countries have built freely for twenty years. Northern Virginia hosts more capacity than most nations; Ireland's grid has been reshaped around it. To insist the footprint is acceptable in Loudoun County but not in Anakapalli requires justification, and “we got there first” is not one. Refusing the Global South the compute the Global North takes for granted, in the name of protecting it, entrenches precisely the dependency that sovereign AI programmes exist to break.

So the case is real. The problem is that almost none of it survives contact with the specifics.

The Arithmetic of Thirst

Water numbers in the data centre debate are unusually slippery, and the slipperiness is not accidental.

The first distinction is between withdrawal and consumption. A facility that withdraws a million litres, passes it through a heat exchanger and returns it a few degrees warmer has consumed almost nothing, though it has still altered the river. A facility that evaporates the same million litres through a cooling tower has consumed all of it. The water is not destroyed, but it has left the basin, and for the people downstream that is the only definition of loss that counts. Headline figures routinely conflate the two.

The second distinction is between on-site and off-site water: what the cooling system uses, versus what was consumed generating the electricity the facility draws. For most of the world's grids, the second dwarfs the first. The third problem is opacity. Nobody outside the companies knows the real figures, because until very recently nobody published them.

Consider what the Visakhapatnam project has declared. According to the News Minute's reporting on the clearance documents, the special purpose vehicles behind the campuses declared a combined water requirement of 501 kilolitres a day, of which 446 would be fresh and 55 recycled. That is roughly half a million litres a day. Set against a city deficit of 70 million litres a day, it sounds trivial.

Now set it against the other numbers in the same documents. The declared power requirement is 1,626 megawatts, plus 971.5 megawatts of diesel backup. A facility drawing more than one and a half gigawatts while consuming half a million litres a day would be, by a wide margin, the most water-efficient large computing installation on Earth. It is not impossible; fully air-cooled designs approach that profile. But it sits uneasily beside estimates from the other direction. Raja Rama Mohan Roy, founder of the non-profit Green Visakha and author of a widely circulated analysis in Countercurrents, puts the cluster's likely freshwater requirement at 55 to 70 million litres a day.

The gap between half a million litres and seventy million is not a rounding error. It is two orders of magnitude, and it exists because there is no mandatory, audited disclosure of data centre water use in India, or almost anywhere else. Mongabay India found that of the fifteen Indian states with dedicated data centre policies, most set no performance standards at all for power or water usage effectiveness, and that a national policy drafted by the electronics and IT ministry has sat unfinalised since 2020. A community cannot argue about a number nobody is required to produce.

The Water You Cannot See

Even if the on-site figure turns out to be genuinely small, the off-site figure will not be.

India's electricity still comes predominantly from coal, and thermal generation is intensely water-hungry because steam turbines require cooling. Under India's Environment (Protection) Amendment Rules, plants commissioned after 1 January 2017 may consume up to three cubic metres of water per megawatt-hour, a limit itself diluted from a stricter 2.5, and one the Centre for Science and Environment has repeatedly documented the sector failing to meet.

Do the arithmetic. A facility drawing one gigawatt continuously for a year consumes roughly 8.76 terawatt-hours. At three cubic metres per megawatt-hour, the embedded water in that electricity is on the order of 26 million cubic metres a year, around 72 million litres a day. That is a back-of-envelope figure, and the real number depends on the generation mix, on how much load is genuinely matched to wind and solar, and on whether the thermal plants use once-through or closed-cycle cooling. But the order of magnitude is the point. The water embedded in the electricity is plausibly a hundred times the water declared for the cooling system.

Google's commitment to build its own renewable generation matters, and in the right direction: solar and wind consume almost no water in operation. But grid-scale matching is an accounting exercise, not a physical one. A data centre running at three in the morning draws electrons from whatever is spinning, and in Andhra Pradesh that is largely coal. The declared diesel backup is a further reminder that keeping a gigawatt of silicon alive on an imperfect grid does not resemble the press release.

This is the analytical heart of the question, and it appears in neither the sustainability report nor the protest banner. The relevant footprint is the pipe entering the building, plus the river cooling the power station, plus the reservoir behind the pumped storage plant that firms the renewables. Count only the first and a gigawatt of compute consumes less water than a mid-sized hotel.

Air Cooling and Its Bill

Google's stated answer is air cooling, and this is a material concession rather than a public relations gesture.

On 3 June 2026 the company published water stewardship commitments that go further than any comparable operator. Google said it would only consider water cooling where local resources are healthy and resilient, and that where a source is at high risk it would choose air cooling or recycled water. It committed to replenishing more water than it consumes across its sites by 2030, a broadening of the 120 per cent replenishment target it had previously set, and put figures behind it: 165 water stewardship projects across 97 watersheds, expected to replenish more than 19 billion gallons a year, more than double what the company consumed in 2024, supported by more than 500 million dollars for water and reuse infrastructure. It also committed to reporting data centre water use by location, the first major cloud provider to do so. Its 2026 environmental report, published on 30 June, stated the company replenished approximately 7.7 billion gallons in 2025, roughly 78 per cent of its freshwater consumption. Google has cited a new Indian data centre as a site where watershed assessment pushed it towards air cooling, and told Reuters the Visakhapatnam project would use advanced air cooling to protect vital local water resources.

Take that at face value and it is a win for the campaigners, and it is also the direct result of pressure. Google did not adopt a watershed risk screen in a vacuum. It adopted one after Querétaro, after Santiago, after years of local resistance in exactly the places where communities had the standing to make trouble.

But air cooling is not free, and the bill lands on the other side of the ledger. Removing heat with air rather than evaporating water requires more fans and airflow, typically around ten per cent more electricity for equivalent cooling. On a coal-heavy grid, that means roughly ten per cent more embedded water consumed at the power station, plus the associated emissions. The net effect can be to move water consumption from one basin to another rather than eliminate it. It is a real improvement for Mudasarlova. Whether it is one for Andhra Pradesh depends on arithmetic nobody has published.

Replenishment carries a similar asterisk. Restoring a wetland in one watershed does not put water back into a different one, and a pledge to return more water than the company consumes is portfolio-wide arithmetic. Nineteen billion gallons replenished across ninety-seven watersheds is a global ratio, and a global ratio says nothing whatever about any particular basin, including the one above Mudasarlova. The commitment means something only if it is basin-specific, same-year and independently verified, and it has not yet been tested in Visakhapatnam.

Nine Days in April

If the water numbers are contested, the process that approved them is not. It is documented, and it is remarkable.

According to the News Minute's account of the clearance record, two special purpose vehicles owned by Adani Infra applied online for environmental clearance on 9 April 2026. The State Level Expert Appraisal Committee met on 10 April. Clearances were granted on 18 April. The foundation stone was laid on 28 April. Nine days elapsed between application and approval for one of the largest industrial developments in the state's history.

The mechanism was a classification decision. The projects were appraised as Category B2 building and construction projects under Schedule 8(a) of the Environment Impact Assessment Notification of 2006, a category requiring neither union government appraisal nor a public hearing. Schedule 8(a) defines it as covering built-up areas below 1.5 lakh square metres. Each of the two Visakhapatnam parks declared a built-up area of roughly 149,000 square metres. E.A.S. Sarma, a retired Indian Administrative Service officer who served as secretary to the Government of India in the power ministry and who lives in Visakhapatnam, has written repeatedly to the environment ministry arguing that the clearances were rushed, leaving little time for meaningful appraisal, and that states sometimes use Category B precisely to avoid central scrutiny. He has asked for them to be revoked.

The deeper problem is structural, and the government has confirmed it in Parliament. Kirti Vardhan Singh, minister of state for environment, forest and climate change, told the Rajya Sabha on 2 April 2026 that AI data centres do not, as such, require environmental clearance under the 2006 notification, and the government restated the position in Parliament in August 2026 with the thresholds spelled out. A data centre needs clearance only if it forms part of a building and construction project with a built-up area above 20,000 square metres, or a township or area development project covering fifty hectares or more, or one with a built-up area of 150,000 square metres or more. India's assessment regime has no category for a facility whose defining impacts are electricity demand, water demand and heat rejection. It regulates the shed, not the machine inside it. A structure is assessed on its square metreage; the gigawatt it draws and the aquifer it taps are, formally, somebody else's department.

The consequences are predictable. The Human Rights Forum has objected that the forest department granted a no objection certificate for a 160-acre site adjacent to Kambalakonda and its notified eco-sensitive zone, and that nearly 90 per cent of the Tarluvada site overlaps the Pedda Chukka Konda reserve forest. Sarma argues that construction is blocking natural water inflows into Mudasarlova. The Forum has filed three cases at the National Green Tribunal. Bolisetti Satyanarayana, national convener of the water conservation network Jal Biradari, has a public interest litigation before the Andhra Pradesh High Court asking it to determine whether statutory environmental and wildlife clearances were required at all for data centre projects near Kambalakonda, and to protect the natural streams and drinking water sources below them. That case was listed for 2 September. The High Court advanced it to 24 August after an application alleging that land levelling, tree felling and other construction activity were continuing on the slopes of the Simhachalam hill range, and directed the respondents to file their answers before the earlier date. A court moving its own timetable forward because the building is outrunning the case is a precise measure of the difficulty: once the participatory channels are closed, litigation is the only instrument left, and litigation moves more slowly than an earthmover. The state denies the project was fast-tracked, and says no water intended for rural or residential use will be diverted, and that the nearby reservoir will not be drawn upon.

Notice what is missing. At no point was there a statutory public hearing. The people of Adavivaram, Tarluvada and Rambilli were not asked. They are litigating because litigation was the only channel left open to them.

The Sovereignty Trilemma

Visakhapatnam is a specific place with specific hills and a specific reservoir. It is also an instance of a pattern that has now been mapped.

In July 2026 a team led by Muntaser Syed, with Marius C. Silaghi, Sheikh Abujar, Sharun Akter Khushbu and Amal El Ahmad, published a study on arXiv titled “The Environmental Cost of Digital Sovereignty: Water, Energy, and Emissions Impacts of Sovereign AI Infrastructure in the Global South”. Nations across the Global South have committed over 200 billion dollars to sovereign AI development, it notes, and the environmental consequences have gone almost entirely unexamined.

The headline finding is stark. Of 52 developing countries with active sovereign AI programmes, roughly 70.6 per cent face high overall water risk according to the World Resources Institute's Aqueduct atlas. The overlap between where sovereign compute is being built and where water is already scarce is close to total. The paper models four cases: the United Arab Emirates, Bangladesh, India and Kenya. A hypothetical 1,024-GPU cluster using evaporative cooling in the UAE would require over 30 million litres a year. Bangladesh's plans, the authors found, contain no siting strategy at all, in a country where more than a fifth of the land floods in an average year and as much as 70 per cent of it in an extreme one.

The authors name the bind precisely, calling it a sovereignty-sustainability trilemma: the difficulty of simultaneously achieving AI sovereignty, environmental responsibility and affordable resources for citizens. Pick any two. A country can have sovereign compute and cheap water if it is prepared to wreck its basins. It can have environmental responsibility and affordable resources if it is prepared to rent its compute from Virginia.

India's position sharpens the point. The World Resources Institute ranks India thirteenth among countries facing extremely high baseline water stress, a group in which agriculture, industry and municipalities withdraw more than 80 per cent of available supply in an average year. This is the country in which the world's largest overseas AI campus is being built.

The global demand curve makes local decisions cumulative. Pengfei Li, Jianyi Yang, Mohammad A. Islam and Shaolei Ren, in the paper that first forced water into the AI conversation, projected that global AI demand would account for 4.2 to 6.6 billion cubic metres of water withdrawal in 2027. The error bars are wide, and its authors say so. But no plausible version of it is small, and every cubic metre lands somewhere specific. Pedram Bakhtiarifard, Pınar Tözün, Christian Igel and Raghavendra Selvan, in a position paper accepted for ICML 2026, argue that treating sustainability as an emissions question alone obscures the tension between expanding access and expanding resource use.

Who Gets To Be Consulted

In July 2026, at the annual session of the United Nations Expert Mechanism on the Rights of Indigenous Peoples, data centres appeared substantially on the agenda for the first time.

Maren Storslett, a member of the Sámi Parliament in Norway, told the meeting that AI is resource-intensive and requires vast amounts of energy, and that in Sápmi large data centres already put immense pressure on their territories. She added a formulation since widely quoted: that the world must not only ask what AI can do, but what it should do, and that respect for the rights of Indigenous peoples must apply. Cheyenna Morgan, an enrolled member of the Keetoowah Band of Cherokee and coalition coordinator for Stop Data Colonialism, put it plainest: these impacts will be felt by regular people who did not ask to have these facilities in their neighbourhoods.

The demand was specific. Delegates called for data centre projects to comply with free, prior and informed consent, for in-depth impact studies before permitting rather than after, and for participation across a project's whole lifecycle. Roberto Anacé, a leader of the Anacé people of northeastern Brazil, described how a ten billion dollar facility proposed near his community had divided it; his people filed a formal complaint in 2025 alleging their consultation rights had been violated. Delegates cited precedents where consent regimes had bitten: a Google data centre in Santiago suspended by a Chilean environmental tribunal in 2024 over an inadequate impact assessment, construction moratoriums adopted by the Seminole Nation of Oklahoma and the Eastern Band of Cherokee Indians, and a 650-megawatt project in Alberta in which the Woodland Cree First Nation holds a 51 per cent stake.

Transpose that framework to Andhra Pradesh and it stops working. India does not accept that the international category of Indigenous peoples applies within its borders. It has not ratified ILO Convention 169, and it voted for the UN Declaration on the Rights of Indigenous Peoples while maintaining that all Indians are indigenous and the declaration is therefore not applicable domestically. What India has instead is a constitutional category, Scheduled Tribes, and a domestic architecture: the Fifth Schedule, designating Scheduled Areas; the Panchayats (Extension to Scheduled Areas) Act of 1996, requiring consultation with the gram sabha before land there is acquired; and the Forest Rights Act of 2006, requiring gram sabha consent for diverting forest land. On paper this is not weaker than free, prior and informed consent. In practice it is chronically under-implemented: Andhra Pradesh published PESA rules in 2011, fifteen years after the Act, and awareness among the communities it protects remains minimal.

Here is what gets lost in the framing. The three campuses are not in Scheduled Areas. Adavivaram, Tarluvada and Rambilli sit in the coastal belt, not the Fifth Schedule agency tracts of the hills, and the people fighting the project in the High Court are urban environmentalists, retired civil servants and residents worried about their taps. Presenting Visakhapatnam as a straightforward story of Indigenous dispossession would be inaccurate. The Adivasi dimension is real, but it is one step upstream, in the electricity.

The Power Comes From the Hills

A gigawatt-scale campus needs firm power, and firm power in a renewables-heavy system needs storage. Andhra Pradesh's answer has been pumped storage hydro, and its pumped storage projects are overwhelmingly located in the Fifth Schedule agency areas of the north of the state.

Reporting by The Wire documented five such projects allocated to private developers, with a combined capacity of 6,600 megawatts across around 2,260 acres in Parvathipuram Manyam and Alluri Sitharama Raju districts, ranging from Karrivalasa at 1,000 megawatts to Pedakota at 1,800. G. Rohit, state secretary of the Human Rights Forum in Andhra Pradesh, said the organisation was shocked at the brazen manner in which the projects had been granted in open contempt of the law, and that no information had been conveyed, no discussion had taken place and there had been no transparency. Ramarao Dora, convenor of the Andhra Pradesh Adivasi Joint Action Committee, said Adivasis in the scheduled areas oppose the projects because they rightly perceive them as harmful. The alleged breaches are specific: Section 5 of PESA, Section 6 of the Forest Rights Act, and the absence of the statutorily required consultation with the Tribal Advisory Council.

These projects predate the Google announcement and are not formally attached to it. That is exactly why they matter. The energy system that will make gigawatt-scale compute viable on India's east coast is being assembled in Adivasi territory, under a consent regime the communities concerned say is being ignored, and no clearance for a campus in Anakapalli will ever ask about it. The impact assessment stops at the fence line. The supply chain does not.

This is the governance gap in its purest form. Consent, where it applies at all, applies to the project on the land. The AI build-out is not a project on a piece of land. It is a system: a campus, a substation, a transmission corridor, a reservoir in the hills, a coal plant burning through a river's cooling capacity. Consent granted or withheld at any single node cannot govern the whole.

Counting the Jobs

The economic case deserves the same scrutiny, and it does not emerge unscathed.

The most careful recent evidence comes from Dany Bahar and Greg Wright, whose analysis of data centre employment effects was published by Brookings in May 2026 and updated on 10 August. Their finding is two-sided. Labour markets receiving their first large data centre see employment in data processing rise by 56 per cent over the first decade, with telecommunications gains of 43 per cent, an effect larger for hyperscale facilities than for colocation providers. But wages remain flat, home prices rise by two to five per cent, and the facilities themselves generate roughly 100 to 200 local jobs each. The authors' summary is that data centres do create jobs, but fewer than industry advocates claim.

Hold that against 188,000. The figure is an ecosystem projection covering construction, indirect and induced employment across a multi-year build-out, not a headcount of people who will work at the campus. Construction employment is genuine, substantial and temporary. Permanent operational staffing is measured in hundreds, because the facilities are automated by design.

That does not make the investment worthless. Subsea cable landings, transmission upgrades and low-latency compute are genuine public goods, and hyperscalers attract other hyperscalers. But it changes how the trade should be evaluated. If the state is discounting land, waiving stamp duty, cutting electricity tariffs for fifteen years and reimbursing tax up to 2,245 crore rupees, the public is buying something and should be told accurately what. Guaranteeing water for twenty years in a city with a 70 million litre daily deficit is itself a subsidy, and it is the one that has not been priced.

What Protection Would Actually Look Like

The gap between what communities in Visakhapatnam can demand and what communities in Santiago or Oklahoma have secured is not one of moral entitlement. It is a gap in legal machinery, and machinery can be built. Six changes would close most of it.

First, make data centres a distinct category under India's Environment Impact Assessment Notification, triggered by connected load and water demand rather than floor area, with mandatory public hearings above a threshold. A facility drawing 1,626 megawatts should not be appraised as a building. The government's own answer to Parliament establishes that it currently is.

Second, require cumulative, basin-level assessment. Visakhapatnam is not receiving one data centre but a cluster, with other operators also planning capacity in the region. Assessing each special purpose vehicle in isolation guarantees the aggregate impact is never examined by anyone.

Third, mandate full water accounting, disclosed publicly per site, independently audited, covering withdrawal and consumption separately and including the embedded water of purchased electricity. Google's per-location reporting is the right template and should be a licence condition rather than a voluntary pledge, applied to every operator. A two-order-of-magnitude gap between declared and estimated consumption should not be possible in a regulated industry.

Fourth, extend consent obligations along the supply chain. If a data centre's firm power depends on pumped storage in a Fifth Schedule area, the gram sabha consent requirements of PESA and the Forest Rights Act should attach to the offtaker as well as the generator. This is not exotic; it is the logic already governing due diligence for conflict minerals and forced labour.

Fifth, price and contract the water properly. Twenty-year guaranteed supply in a deficit basin should carry a full-cost tariff, a hard volumetric cap, automatic curtailment ahead of domestic and agricultural users during declared scarcity, and same-basin replenishment verified annually by an independent body. Credits earned in another watershed should not count.

Sixth, negotiate a binding community benefit agreement rather than relying on corporate social responsibility. The Alberta precedent, where a First Nation holds a majority stake in a 650-megawatt facility, shows the alternative to exclusion is not obstruction but ownership. Equity, local hiring guarantees, funded municipal water infrastructure and an enforceable grievance mechanism turn a community from an obstacle into a counterparty.

None of this requires India to accept a category of Indigenous peoples it has rejected for four decades. Every item is achievable inside its existing constitutional framework. They require only that the framework be applied.

Answering the Question From the Hillside

So what does it mean when the technology powering artificial intelligence in wealthier nations is fuelled by water drawn from a set of countries composed, almost entirely, of countries that cannot spare it? Of the 52 developing nations with active sovereign AI programmes, 70.6 per cent already face high water risk. That is not a claim about the developing world in general. It is a claim about the particular list of countries that have decided to build their own compute, and that list is made up, very nearly without exception, of places where the water has already run short.

It means, first, that the transaction is not what it appears. The exchange offered to Visakhapatnam is not water for prosperity. It is water, land, forest, forgone revenue, discounted power and twenty years of guaranteed supply, in return for a few hundred permanent jobs, a modest ecosystem effect, and a claim on digital sovereignty that will be exercised mainly by a company headquartered in Mountain View. That may still be a trade worth making. It is not the trade that was described.

It means, second, that the injustice is procedural before it is environmental. Nine days is not enough time to appraise anything, and a classification that avoids a public hearing is not a technicality; it is the whole ballgame. The people of Adavivaram did not lose an argument about water. They were never given the argument. Everything that has followed, the marches, the banners, the tribunal cases, the litigation the High Court pulled forward to keep pace with the diggers, is the sound of a community using the only instruments left after the participatory ones were bypassed.

It means, third, that the exported thirst is invisible even to the people exporting it. The user in London generating a summary cannot know that the marginal electron came from a coal plant in Andhra Pradesh, that the plant consumed three cubic metres of water per megawatt-hour, or that the reservoir firming the renewables sits on land an Adivasi gram sabha was never asked about. The trilemma identified by Syed and colleagues is invisible from the prompt box by design.

And what protections should exist? Not a veto exercised from abroad. The paternalism objection is right that the Global South should not be denied infrastructure the Global North built without asking anyone. But it cuts the other way, and harder. Communities in Chile, in Oklahoma and in Ireland have won moratoriums, suspensions, disclosure requirements and equity stakes, because they had standing, hearings, tribunals that would listen, and in some cases treaty rights. The paternalistic act is not insisting that Visakhapatnam gets protections. It is building at Visakhapatnam precisely because it will not.

The protections that should exist are the ones that already exist elsewhere: a statutory right to be heard before the earth is cut, audited disclosure of what will be consumed, consent obligations that follow the electricity to its source, and a share of the equity rather than a plaque on a wall. There is nothing in that list Google could not accept tomorrow and still make money, and nothing Andhra Pradesh lacks the power to require.

Work on the hillside above Mudasarlova continues. The terraces are cut, the red earth is exposed, and the monsoon runoff that used to feed the reservoir now finds a different path down. The High Court's hearing is listed for this week, the matter live before it as this goes to press. Whatever it decides, and whenever, the more consequential question was settled by default in nine days in April, by a form that classified a gigawatt of artificial intelligence as a building. The water will follow the law. The only thing still undecided is whose law it is.

Sources and References

  1. Reuters, “Google's $15 billion India data centre project battles water, wildlife concerns”, 6 August 2026, syndicated by The Kathmandu Post. https://kathmandupost.com/world/2026/08/06/google-s-15-billion-india-data-centre-project-battles-water-wildlife-concerns
  2. edie, “Google's $15bn India data centre project faces water and environmental challenges”, August 2026. https://www.edie.net/googles-15bn-india-data-centre-project-faces-water-and-environmental-challenges/
  3. Google, “Our first AI hub in India, powered by a $15 billion investment”, Google India Blog, 14 October 2025. https://blog.google/intl/en-in/company-news/our-first-ai-hub-in-india-powered-by-a-15-billion-investment/
  4. The Tribune (India), “Andhra govt allots 480 acres for Adani-Google AI data Centre in Visakhapatnam”. https://www.tribuneindia.com/news/business/andhra-govt-allots-480-acres-for-adani-google-ai-data-centre-in-visakhapatnam/
  5. Data Center Dynamics, “Google confirms $15bn data center project in Andhra Pradesh, India”. https://www.datacenterdynamics.com/en/news/google-confirms-15bn-data-center-project-in-andhra-pradesh-india/
  6. The News Minute, “Inside the subsidies for Google's Vizag data centre amid Andhra-Karnataka spar”. https://www.thenewsminute.com/andhra-pradesh/inside-the-subsidies-for-googles-vizag-data-centre-amid-andhra-karnataka-spar
  7. The News Minute, “Vizag Google Adani AI data centre: State govt rushed environmental clearance, activists allege”. https://www.thenewsminute.com/andhra-pradesh/vizag-google-adani-ai-data-centre-state-govt-rushed-environmental-clearance-activists-allege
  8. The News Minute, “HRF challenges legality of Adani-Google Vizag data centre, demands immediate stop”. https://www.thenewsminute.com/andhra-pradesh/hrf-challenges-legality-of-adani-google-vizag-data-centre-demands-immediate-stop
  9. Mongabay India, “India bets on data centres even as water, energy-use concerns mount”, May 2026. https://india.mongabay.com/2026/05/india-bets-on-data-centres-even-as-water-energy-use-concerns-mount/
  10. Mongabay, “Indigenous advocates push for rights protections around AI data centers”, July 2026. https://news.mongabay.com/2026/07/indigenous-advocates-push-for-rights-protections-around-ai-data-centers/
  11. Business & Human Rights Resource Centre, “Global: Indigenous leaders call for data centres projects to comply with their right to free, prior and informed consent during UN meeting”, July 2026. https://www.business-humanrights.org/en/latest-news/global-indigenous-leaders-call-for-data-centres-projects-to-comply-with-their-right-to-free-prior-and-informed-consent-during-un-meeting/
  12. Raja Rama Mohan Roy, “The Hidden Environmental Cost of AI: Can Visakhapatnam Afford Giant Data Centres?“, Countercurrents, July 2026. https://countercurrents.org/2026/07/the-hidden-environmental-cost-of-ai-can-visakhapatnam-afford-giant-data-centres/
  13. Countercurrents, “AI Data Centres and Environmental Justice: E.A.S. Sarma Urges Scrutiny of Projects Amid Ecological and Human Rights Concerns”, June 2026. https://countercurrents.org/2026/06/ai-data-centres-and-environmental-justice-e-a-s-sarma-urges-scrutiny-of-projects-amid-ecological-and-human-rights-concerns/
  14. Countercurrents, “HRF Demands Halt to Vizag Hyperscale Data Center Over Alleged Environmental Violations”, May 2026. https://countercurrents.org/2026/05/hrf-demands-halt-to-vizag-hyperscale-data-center-over-alleged-environmental-violations/
  15. Deccan Chronicle, “High Court Advances Hearing Over Data Centres Near Kambalakonda Sanctuary”. https://www.deccanchronicle.com/southern-states/andhra-pradesh/high-court-advances-hearing-over-data-centres-near-kambalakonda-sanctuary-1976576
  16. Muntaser Syed, Marius C. Silaghi, Sheikh Abujar, Sharun Akter Khushbu and Amal El Ahmad, “The Environmental Cost of Digital Sovereignty: Water, Energy, and Emissions Impacts of Sovereign AI Infrastructure in the Global South”, arXiv:2607.13443, 15 July 2026. https://arxiv.org/abs/2607.13443
  17. Pengfei Li, Jianyi Yang, Mohammad A. Islam and Shaolei Ren, “Making AI Less 'Thirsty': Uncovering and Addressing the Secret Water Footprint of AI Models”, arXiv:2304.03271. https://arxiv.org/abs/2304.03271
  18. Pedram Bakhtiarifard, Pınar Tözün, Christian Igel and Raghavendra Selvan, “Position: Neglecting the Sustainability of AI is Fuelling a Global AI Arms Race”, arXiv:2502.20016, accepted at ICML 2026. https://arxiv.org/abs/2502.20016
  19. World Resources Institute, “25 Countries, Housing One-quarter of the Population, Face Extremely High Water Stress”, Aqueduct Water Risk Atlas. https://www.wri.org/insights/highest-water-stressed-countries
  20. Google, “Google's water stewardship commitments for local communities”, 3 June 2026. https://blog.google/company-news/outreach-and-initiatives/sustainability/new-water-stewardship-commitments/
  21. Google, “Read Google's 2026 Environmental Report”, 30 June 2026. https://blog.google/company-news/outreach-and-initiatives/sustainability/2026-environmental-report/
  22. Down To Earth, “As told to Parliament (April 2, 2026): Centre acknowledges AI's energy and water footprint while key facilities remain beyond EIA mandate”. https://www.downtoearth.org.in/environment/as-told-to-parliament-april-2-2026-centre-acknowledges-ais-energy-and-water-footprint-while-key-facilities-remain-beyond-eia-mandate
  23. The Print, “AI data centres do not need environment clearance, Centre tells Parliament”, August 2026. https://theprint.in/environment/ai-data-centres-environment-clearance-centre-parliament/3007640/
  24. The Wire, “Andhra Pradesh: Pumped Storage Projects Spark Concerns over Tribal Displacement and Environmental Harm”, 6 December 2024. https://m.thewire.in/article/environment/andhra-pradesh-pumped-storage-projects-spark-concerns-over-tribal-displacement-and-environmental-harm
  25. Dany Bahar and Greg Wright, “New evidence on data center employment effects”, Brookings Institution, published 4 May 2026, updated 10 August 2026. https://www.brookings.edu/articles/new-evidence-on-data-center-employment-effects/

Tim Green

Tim Green UK-based Systems Theorist & Independent Technology Writer

Tim explores the intersections of artificial intelligence, decentralised cognition, and posthuman ethics. His work, published at smarterarticles.co.uk, challenges dominant narratives of technological progress while proposing interdisciplinary frameworks for collective intelligence and digital stewardship.

His writing has been featured on Ground News and shared by independent researchers across both academic and technological communities.

ORCID: 0009-0002-0156-9795 Email: tim@smarterarticles.co.uk

Listen to the free weekly SmarterArticles Podcast

Discuss...

Somewhere tonight, a person will leave a hospital holding a sheet of paper they can read perfectly and should not trust.

The paper will look right. Clean typography, the hospital's logo in the corner, sentences running in grammatical order in the reader's own language. It will have been produced in under a second at a cost that rounds to nothing, and handed over by a clinician who cannot read a word of it, to a patient with no way of comparing it to anything. Everyone will believe the document is complete, because it looks complete, and looking complete is what this technology does flawlessly.

We know what these documents can contain, because researchers have checked. In 2021, a team led by Breena Taira at the University of California, Los Angeles took twenty commonly used emergency department discharge statements, ran them through Google Translate into seven widely spoken languages, and had native speakers evaluate the output. Their paper in the Journal of General Internal Medicine contains two findings that ought to have ended the argument about unsupervised machine translation in clinical settings. The first is statistical. The second is a pair of words: in the Armenian output, ibuprofen was rendered as “anti-tank missile”, and in Chinese, the anticoagulant Coumadin came out as “soybean”.

These are not subtle failures. A bilingual reader would spot them in a second. The problem is that no bilingual reader was ever going to look, because the entire economic point of the exercise was to avoid paying one.

That is the shape of the thing. Over roughly three years, the cost of producing a translation has fallen by close to three orders of magnitude. The cost of finding out whether it is correct has not moved, because verification still requires a person who reads both languages, sitting down with both documents, and thinking. Generation collapsed. Verification did not. Everything that follows, from the restructuring of a seventy billion dollar industry to the question of who is answerable when the instructions are wrong, falls out of that single asymmetry.

The Arithmetic That Broke in Only One Direction

Google's Cloud Translation service prices standard neural machine translation at twenty United States dollars per million characters, the first half a million each month free. DeepL's professional interface has been offered at around five and a half dollars per million. A thousand words of English is roughly six thousand characters, so translating it costs about twelve cents at the top rate and about three at the bottom. Professional human translation of the same thousand words sits, in 2026, at one hundred and fifty to three hundred dollars, with rate surveys clustering human work between ten and thirty cents per source word.

Generation is effectively free. But consider what verification requires and the picture inverts. To know whether a thousand-word translation is correct, a qualified bilingual reader must read the source, read the target, and compare them for meaning, register, omission, addition and negation. That is not faster than translating. It is often slower, because the reviewer must reconstruct the translator's reasoning as well as the author's. And it requires precisely the scarce human capability the machine was bought to replace.

Verification now costs more than production by a factor of roughly a thousand. If a factory could stamp a component for a tenth of a cent but testing each one cost a hundred dollars, nobody would call the testing optional. They would redesign the process, or admit on the record that they were shipping untested parts. Translation took the third option, never written down: shipping untested parts while behaving as though they had been tested, because they look exactly like the tested ones.

A second scarcity sits beneath the first. The binding constraint is not money but qualified bilingual attention. Only so many people can competently review a discharge instruction in Armenian, a tenancy notice in Tigrinya, or a pesticide warning in Hmong. Machine translation multiplied the volume of text requiring review by orders of magnitude while leaving the pool of reviewers where it was. The bottleneck is a population, not a budget line.

Failure With No Symptoms

The asymmetry persists because its consequences are invisible at the point of delivery. Machine translation does not fail the way software fails. It does not crash, return an error code, garble characters or leave a field blank. It fails by producing a well-formed sentence that means something other than what the source said.

The canonical example comes from a 2019 study in JAMA Internal Medicine by Elaine Khoong, Eric Steinbrook, Cortlyn Brown and Alicia Fernandez, who passed one hundred sets of emergency department discharge instructions comprising 647 sentences through Google Translate into Spanish and Chinese. Eight per cent of the Spanish sentences and nineteen per cent of the Chinese sentences were inaccurate. More to the point, two per cent of the Spanish and eight per cent of the Chinese carried potential to cause clinically significant harm.

One sentence read, in English, “hold the kidney medicine until you have a chance to speak with your kidney doctor”. The Chinese output instructed the patient to keep taking the kidney medicine, and the Spanish said much the same. “Hold”, in clinical English, means stop. The machine took it in its ordinary sense of retain and produced a fluent, grammatically flawless instruction to do the opposite of what the prescriber intended. Elsewhere, “aortic aneurysm” became, in Spanish, an “evacuation of the main blood vessel”.

The pattern is not confined to hospitals. A 2010 study in Pediatrics by Iman Sharif and Julia Tse examined 76 Spanish-language prescription labels generated by the computer programmes New York pharmacies were actually using and found errors in half of them, including the notorious collision in which the English word “once” reads in Spanish as eleven. In immigration proceedings, translators with the volunteer organisation Respond Crisis Translation told Rest of World in 2023 that machine translation of Pashto and Dari was corrupting asylum claims, in one case rendering the first person singular as a plural throughout a personal narrative, so that an individual persecution story read as a collective complaint. At least one Afghan claim was rejected after such errors.

No surface signal attaches to any of these. A negation reversal reads exactly like a negation preserved: same length, same cadence, same register, same institutional wrapper. That is what distinguishes machine translation from almost every other failure mode in consumer technology. The output carries no evidence of its own unreliability, and the evidence that would reveal it exists only in a document the recipient does not have.

Consider the chain. The clinician cannot read the target language and assumes the system did its job. The organisation that licensed the engine sees an aggregate quality score, if anything. The patient sees fluent prose on hospital letterhead and assumes that if it had not been checked, it would not have been handed over. Nobody holds both halves of the comparison, and the appearance of completeness does all the work verification used to do.

The older failure mode is instructive by contrast. In January 1980, an eighteen-year-old named Willie Ramirez arrived unconscious at a Florida emergency department. His Cuban family said he was “intoxicado”, meaning ill from something he had eaten. It was taken to mean intoxicated. Ramirez was treated as a drug overdose, his intracerebellar haemorrhage went undiagnosed for two days, and he was left quadriplegic; the settlement was estimated at around seventy-one million dollars over his lifetime. That failure had a location: a room, a decision, a defendant. When the same error comes from an engine invoked automatically inside an electronic patient record, it has no location at all.

A Thirty-Nine Point Gap Inside One Product

The Taira study's statistical finding should trouble anyone who believes machine translation is solved. Across seven languages, the overall meaning survived in 82.5 per cent of cases, 330 out of 400. That headline is meaningless, because the variation underneath it is enormous.

Spanish came out at 94 per cent accuracy. Tagalog at 90. Korean at 82.5. Chinese at 81.7. Vietnamese at 77.5. Farsi at 67.5. Armenian at 55. A thirty-nine point spread between best and worst inside a single product, through a single interface, with a single set of expectations attached. Farsi had a further problem no accuracy metric captures: directional text handling rendered some of it illegible.

Nothing in the interface communicates that spread. The menu presents Armenian and Spanish as equivalent options, and output arrives with the same speed, formatting and absence of caveat. An administrator enabling automated discharge translation is not choosing between a reliable service and an unreliable one. They are enabling both at once, with no signal for which patients get which.

This is language-tier inequality, and it maps onto each language's digital footprint. Languages with vast parallel corpora perform well; those spoken largely by populations who were never a lucrative localisation market perform badly. So the patients most likely to receive an unreliable translation are those least likely to have an alternative: recent arrivals, smaller diaspora communities, speakers of languages for which the local health system has no on-call interpreter at three in the morning.

The gap is not closing, because progress concentrates where the data is. A study in JMIR Formative Research in January 2026 by a team at Beth Israel Deaconess Medical Center tested Claude Sonnet 3.5 on Spanish translations of emergency discharge instructions and found the output essentially clinically acceptable, with evaluators scoring completeness and severity at 5.0 on a five-point scale and meaning at 4.9. Genuinely impressive, and a result about Spanish. It is also no longer isolated, which is the part of this argument that has changed. Meanwhile a framework paper in npj Digital Medicine in September 2025 by Ivan Lopez, David Velasquez, Jonathan Chen and Jorge Rodriguez recommends deploying first into languages with abundant digital resources, naming Spanish and Portuguese, and warns that underrepresented languages such as Quechua and Yorùbá need further testing first. The authors are right. But notice what that implies: the languages validated first needed validating least, and those where risk is highest get served last, or with no validation at all.

The strongest evidence against the case being made here arrived in Academic Emergency Medicine in 2026, and it deserves stating at full strength rather than filing under limitations. Giovanni Rodriguez, Patricia Hernández and colleagues ran a blinded non-inferiority study on fifty-three real emergency department discharge instructions of between one hundred and five hundred words, taken exactly as the clinicians wrote them, original spelling and grammar errors left in. Google Translate and ChatGPT-4o rendered them into Spanish, Brazilian Portuguese and Simplified Chinese, against professional translations of the same text. Professional medical interpreters, blinded to which was which, scored 477 unique translations for fluency, adequacy, meaning and severity against a non-inferiority margin of half a point.

Both machines came out non-inferior to the professionals on most domains. In Spanish and Brazilian Portuguese they matched professional translation on adequacy, meaning and severity, falling short only on fluency. In Simplified Chinese they matched it on all four. The frequency of clinically significant errors did not differ significantly by translation method.

That has to be conceded without qualification, because it is true and it matters. In those three languages, against genuinely messy clinician-written source text rather than tidied research prose, current machine translation is not measurably more dangerous than a professional human being. The 2019 and 2021 numbers quoted earlier are real, and they were real about the systems they tested, but those systems are two generations back. Nobody should now cite eight per cent Spanish inaccuracy as a description of what a modern engine does to a Spanish discharge sheet. It is not what happens.

What the study does not do is rescue the position it appears to demolish, and both reasons are visible in its own design. Start with the languages. Spanish, Brazilian Portuguese and Simplified Chinese are among the largest localisation markets on earth and among the best-resourced pairs in existence. The study says nothing whatever about Armenian at 55 per cent, or Farsi at 67.5, or any language whose speakers were never worth building a corpus for. It does not narrow the thirty-nine point spread. It measures the top of it, carefully, one study later.

The second reason goes to what non-inferiority establishes. That clinically significant errors did not differ by translation method is not a finding that machine translation is safe. It is a finding that professional human translation also carries a clinically significant error rate, and that the machine has drawn level with it. Nobody who commissions professional translation of consequential documents believes a single pass is sufficient, which is why the profession has review stages, standards describing them, and indemnity behind them. Both arms of that study were then read by trained evaluators. Every translation in it was checked. The subject here is the document nobody checks at all, and non-inferiority to an unverified human baseline is not a warrant for skipping verification. It moves the argument off provenance and onto verification, which is exactly where it belongs. The question was never whether a machine or a person produced the text. It is whether anyone competent read it afterwards.

There is direct evidence on both counts, from a study that measured what the non-inferiority design left out. In npj Digital Medicine in October 2025, Ryan Brewster and a large multidisciplinary team took paediatric discharge instructions into six languages, Arabic, Armenian, Bengali, simplified Chinese, Somali and Spanish, and compared three routes: ChatGPT-4o alone, professional linguist translation, and human-in-the-loop, meaning a machine draft post-edited by a professional linguist. Forty-two evaluators scored the output, twelve professional linguists, sixteen clinicians and fourteen bilingual family caregivers.

The tier gap turned up intact in a frontier model. For Spanish and Bengali, ChatGPT-4o came close to professional quality. For Armenian it scored 2.4 on overall quality against 3.6 for professional translation, a deficit of more than a full point on a five-point scale, and Somali and simplified Chinese also fell significantly below the professional standard. Whatever has been fixed since 2021, this has not been. Note that the two studies disagree about Chinese, non-inferior on every domain in one and materially below professional standard in the other, on different document types, against different comparison standards, judged by different evaluators, a year apart. That disagreement is not a scandal. It is a measurement of how thin the evidence is.

The other finding is the one this whole argument turns on. Human-in-the-loop did not merely rescue Armenian, it beat the professionals, scoring 3.9 against their 3.6. Across all six languages it was also the faster route, averaging 7.1 minutes per document against 16.8 for professional translation. Post-editing a machine draft was better than either route alone and roughly twice as fast as translating from scratch. That is what working looks like, and it is worth being precise about what it costs. It still requires the qualified bilingual reader. It makes that person faster. It does not make them unnecessary.

Every such study is itself an act of verification, and inherits the same cost structure. Beth Israel validated one language pair, one document type, one institution, which tells you nothing about Armenian, consent forms, or next quarter's model update. The evidence base is built at human speed against a deployment surface expanding at machine speed.

The Instruments That Cannot See the Errors That Matter

The obvious response is to automate the checking. If a machine can produce the translation, surely a machine can grade it. That is the promise of quality estimation, and what most enterprise deployments rely on when they route some segments to review and let others through.

Two pieces of 2026 research suggest the hope is misplaced exactly where it matters. In August, Serge Gladkoff, Angelika Vaasa, Sue Ellen Wright, Ingemar Strandvik and Lifeng Han released a peer-reviewed paper titled “Looking under the Wrong Lamppost”, accepted for the ninth International Conference on Natural Language and Speech Processing in September, arguing that automated quality estimation suffers structural rather than incidental limitations. Their conclusion is unusually blunt: segment-level scores should not be used as a standalone basis for routing, release or review bypass in production. They document overfitting, distribution collapse, failure to generalise across domains, and blindness to cohesion and coherence, which are not properties of individual segments at all. Their proposed direction is not better scoring but making human review cheaper rather than pretending it is unnecessary.

The clinical version arrived two months earlier. At the American Medical Informatics Association's Amplify conference in June 2026, William Mundo, Elizabeth Goldberg and Yanjun Gao presented work on AI-generated translation of emergency discharge instructions and found direct disagreement between automated evaluation and clinician judgement. Automated metrics suggested high quality. Clinicians found errors in precisely the meaning-sensitive areas carrying clinical consequence: mistranslated medication instructions, omitted return precautions, negation errors of the “take with food” variety, confusion between daily and twice daily, mild discomfort turned into severe pain. Automated metrics alone, they concluded, should not drive deployment.

Read together, the escape route closes. The class of error automated estimation misses is not a random sample. It is disproportionately the class that reverses instructions, drops warnings and changes doses, because those errors are small, local, fluent and invisible to systems trained to reward fluency. The scoring machinery is best at catching failures a human would catch instantly, and worst at the failures only a human can catch at all. You cannot build a verification layer out of the competence that produced the problem.

What the Market Did With the Savings

To see where value goes when a commodity becomes free, read the 2026 Nimdzi 100, the annual ranking of the world's largest language industry providers. It sizes the market at 72.6 billion dollars and reports growth figures that read like a diagram of stratification.

The hundred largest providers collectively grew by 1.1 per cent. The top ten grew by 3.6 per cent. Providers ranked fifty-first to hundredth contracted by 4.3 per cent, the first negative growth recorded for any segment since 2021. Ten providers entered the ranking who were not there the year before, largely through acquisition; Propio Language Services vaulted into third place by acquiring the interpreting company CyraCom. Industry optimism, on Nimdzi's own confidence rating, fell from 6.5 to 6.2.

That distribution is the signature of a market whose commodity has moved to machines. If you sell volume translation, your product competes with something priced at twelve cents per thousand words, and no cost base survives that. If you sell verification, liability, domain expertise, certification and the ability to put a name and an indemnity policy behind a document, you are selling the one thing machines did not make cheaper. The middle contracts because the middle sold the thing that became free.

From inside the profession, the same restructuring reads differently. A study submitted to arXiv in June 2026 by Yujun Wang, Ehud Reiter, Shimei Pan, Steffen Eger and Wei Zhao analysed 79,286 social media posts from Reddit, Facebook, Bluesky and Mastodon between 2019 and 2025 across four stakeholder groups. The arc in translator discussion tells the story in three data points. In 2021 they were talking about computer-assisted translation infrastructure, translation memories and termbases. By 2023 it was integrating AI into those workflows. By 2025 it was automation and quality assurance, which is to say they had stopped discussing how to translate and started discussing how to check machines. Translators held the most critical stance of any group. The communities are not even arguing about the same variables: where developers frame reliability as a hallucination rate, everyone else frames it as verification and oversight.

The professional consequence is a job description inverted without being renegotiated. Translators are increasingly employed not to produce text but to verify it, the more cognitively demanding half, and are paid at post-editing rates set on the assumption the machine already did most of the job. The standards infrastructure sees the distinction even where the market does not. ISO 17100 covers human translation services; ISO 18587 covers post-editing of machine output and requires post-editors to meet the same competence levels as translators under ISO 17100. Same competence, different pay grade, different assumption about how much of the output anyone read.

Nobody Ever Decided This

The most striking feature of this migration is that almost nowhere can you find the meeting where it was approved.

Organisations did not convene a risk committee, evaluate error rates by language, define a threshold above which human review became mandatory, and sign off. They enabled a feature. A content management system offered automatic localisation; a support platform added a translate button. A free allowance covered the first half a million characters a month, enough to translate a small organisation's entire public documentation set without generating an invoice that needed a signature.

Research commissioned by DeepL and reported in 2026 found enterprises investing heavily in language AI while still running critical global operations through manual, unintegrated workflows, which tells you the adoption pattern is not a strategy but a scattering of local decisions. Marketing copy, support documentation, product listings, internal policy, contract summaries and patient-facing instructions crossed the same threshold at different moments, authorised by different people, none asked to consider the whole. A system nobody decided to build is one nobody feels responsible for.

The Chain in Which Everyone Is a Bystander

Try to identify the responsible party in a concrete case and the exercise becomes a study in evaporation.

A patient is discharged with machine-translated instructions containing a reversed medication direction. Is the model provider liable? Its terms disclaim high-stakes use and it never knew the text was clinical. The record vendor? It integrated a service at a customer's request. The hospital? Arguably yes, the most promising thread, but it will point to the impossibility of reviewing every document in every language, and to the absence of any regulatory instruction telling it where the line sits. The clinician cannot read Armenian. The patient received a document from a hospital in their own language and read it correctly. The document was wrong.

Occasionally an institution refuses to accept this diffusion. In June 2018, a United States district court in Kansas suppressed evidence in United States versus Cruz-Zamora, in which an officer used Google Translate on a patrol car laptop to obtain consent to search a vehicle from a driver who spoke very little English. The officer typed “can I search the car?” The Spanish output, translated back, asked something closer to “can I find the car?” The court found the literal but nonsensical rendering meant the driver had not given unequivocal consent. That is a court locating responsibility with the party who chose to deploy the technology, rather than the person who received its output.

Cruz-Zamora is instructive precisely because it is unusual. It required an adversarial process, a defence lawyer, a back-translation and a judge willing to inspect the mechanism. Almost no machine-translated document gets that treatment. The vast majority arrive where there is no adversary and nobody whose job it is to check: a discharge sheet, a benefits letter, a tenancy notice, a warning label, a consent form. The error surfaces, if ever, only through the harm it causes, by which point the chain back to a mistranslated clause is nearly impossible to reconstruct.

The industry sells professional indemnity insurance precisely because translation carries assignable liability when a professional does it. A named translator working under ISO 17100 is insurable because the responsible party is identifiable. A model invocation inside a content pipeline is not, and the risk does not disappear when it becomes uninsurable. It transfers to the person holding the paper.

The Law Points Two Ways at Once

For a case study in regulatory incoherence, consider the current United States position on language access.

Section 1557 of the Affordable Care Act, whose final implementing rule took effect on 5 July 2024, requires covered health entities to provide qualified interpretation and translation, and specifically requires that machine-translated material be reviewed by a qualified human translator where accuracy is essential or the source is complex or technical. That addresses the verification gap directly: you may use the machine, but not unsupervised where it matters.

Eight months later, on 1 March 2025, Executive Order 14224 designated English the official language of the United States and revoked Executive Order 13166, the 2000 order requiring federal agencies to plan for meaningful access for people with limited English proficiency. It directed the Attorney General to rescind guidance issued under it, and the resulting Department of Justice memorandum of 14 July 2025 encouraged agencies to rescind their own guidance, to consider English-only services, and, explicitly, to use artificial intelligence to reduce translation costs. In July 2026, formal rescission notices for Title VI language access guidance appeared in the Federal Register.

The policy environment now simultaneously requires human review of machine translation in health settings and recommends machine translation as a cost-reduction strategy across federal programmes. Underlying civil rights law has not changed: Title VI still prohibits national origin discrimination, and recipients of federal funding still owe meaningful access. What has changed is the guidance telling them how, so the question of what meaningful access means when the translation is machine-generated is unanswered by anyone with authority to answer it.

Europe's failure has a different texture. The General Product Safety Regulation, applicable since 13 December 2024, requires products to be accompanied by instructions and safety information in a language easily understood by consumers in the member state where they are sold. That is an outcome requirement, usually preferable to a process requirement. But it is silent on how the translation is produced and on who must verify the text is in fact understood rather than merely present. A machine-translated warning that reverses a prohibition satisfies the formal requirement while defeating its purpose. The EU Artificial Intelligence Act adds transparency obligations from August 2026, including machine-readable marking of synthetic content, though whether a translation engine running over patient instructions falls inside its high-risk regime is a matter for lawyers rather than settled fact. A small joke is buried here: some widely consulted online reference versions of the AI Act's own articles carry a notice explaining that the translations are machine-generated and not the official European Parliament versions.

In the United Kingdom, NHS England published an improvement framework for community language translation and interpreting on 27 May 2025 that is unusually honest about the gap. It records annual NHS spend at around 75.5 million pounds against an estimated 250 to 300 million needed to meet actual demand, and warns that translation apps, while convenient, carry risks to accuracy and patient safety. It commits NHS England to national guidance specifying when AI is suitable, which tools are approved, how accuracy is verified across languages, and the governance frameworks including indemnity and responsibility. That last clause is an admission: who is responsible when an AI translation harms an NHS patient had not been settled. It is being written now, after the tools are already on the wards.

What Recipients Are Actually Being Handed

Strip away the institutional framing and answer the first question plainly. What does it mean to receive information machine-translated at a fraction of human cost and never checked by anyone who reads both languages?

It means being handed a document carrying the full social signalling of institutional assurance and none of the substance. Every cue on that page, the letterhead, the formatting, the fluency, evolved in a world where producing text in a language required a person who knew it. Those cues were reliable because they were expensive. Fluency was a costly signal of competence. Machine translation made fluency free, so the signal conveys nothing, and nobody has told the recipients.

That is usually asserted about lay recipients, and it would be easy to dismiss as condescension: of course a patient cannot tell, they are not a translator. But the non-inferiority study tested the claim without setting out to. Its blinded evaluators were professional medical interpreters, reading in their own working languages, paid to assess translation quality, and aware that some of what they were reading was machine output. They frequently took the machine translations for professional work. If the signal has stopped carrying information under those conditions, for those readers, there is no version of this in which it still carries information for a patient reading a discharge sheet in a corridor. Fluency is not weak evidence of human competence. It is no evidence at all.

It means epistemic risk has been transferred to the party least able to bear it. Previously the burden of ensuring correctness sat with the institution producing the text, because the institution employed the translator. It has been silently relocated to the reader, who is expected, without being told, to treat official documents in their own language as provisional. Those readers are disproportionately the people with fewest resources to seek a second opinion.

It means informed consent is being quietly hollowed out. Consent requires comprehension of accurate information. A patient who fully understands a fluent instruction that has reversed the source has comprehended something, but not the thing they were meant to consent to. Their signature attests to a document nobody on the institutional side has read.

And it means we have built, without deciding to, a two-tier information citizenship. Speak a high-resource language and the machine serves you at close to professional quality. Speak Armenian and the same interface, the same button, the same institutional promise delivers something accurate slightly more than half the time in controlled testing. Both are told they have been given information in their own language. Only one of them has.

An Accountability Regime That Would Actually Bite

The second question is harder. Who bears responsibility when the system producing the translation is more capable than any system available to tell whether it worked?

The honest starting point is that no technical fix resolves this. The verification gap is not an artefact of immature models but a structural property of a task whose correctness can only be assessed by someone holding both languages. Better models will narrow the error rate. They will not create the reviewers. So the regime must be built out of allocation of responsibility rather than engineering, starting by reversing the current default.

That default should be simple. The organisation publishing a translated document warrants it, regardless of how it was produced. Not the model provider, not the platform, not the clinician, and emphatically not the recipient. If a hospital hands a patient a discharge sheet in Armenian, it stands behind that sheet exactly as if a staff translator had written it. This single rule does most of the work, because it puts the cost of unverified translation onto the party that captured the savings rather than the recipient who absorbs the risk.

The rest follows. Risk tiering should be organised by consequence class rather than document type. A marketing headline and a dosing instruction are both text; only one can put someone in intensive care. Content that can change a medical action, alter a legal obligation, forfeit a right, or carry a safety warning belongs in a tier requiring independent bilingual review before release, attributable to a named professional. Everything else can run unverified, provided it is labelled as such.

Per-language performance disclosure should be mandatory and visible in the interface. A product performing at 94 per cent in one language and 55 per cent in another should not present those as equivalent menu options. Deployers in regulated settings should publish language-specific error rates for their own content domains, sampled by qualified reviewers, and suspend automated release for pairs below a defined floor. The Lopez framework recommends the healthcare version: extend Joint Commission oversight to machine-assisted translation, build a shared clinical translation corpus, and update Section 1557 guidance with language-specific benchmarks. Language-specific is the operative phrase, because aggregate accuracy conceals the inequality that matters.

Provenance labelling should appear in the target language, in the document, in terms a recipient can act on. Not a disclaimer in eight-point type, but a plain statement: this text was produced automatically and has not been checked by a person who reads both languages; if anything here concerns your medication, contact us. The harm mechanism is a false impression of human authorship, and labelling is the cheapest counterweight.

Procurement is the most immediately available lever. Contracts in regulated settings should specify ISO 18587 full post-editing as a floor for consequence-bearing content, name the accountable reviewer, require logging sufficient to reconstruct which engine produced which document on which date, and prohibit reliance on automated quality estimation as the sole basis for bypassing review. The Gladkoff paper gives that last clause its evidentiary basis. The Brewster results give the first clause its own: full post-editing was not a compromise between speed and safety but better than either route alone, and faster than commissioning a translation from scratch. The expensive part of translation is no longer the translating. It is the reading.

Finally there is a public-goods problem no single organisation will solve. Armenian performs at 55 per cent because of a shortage of high-quality parallel data, and that shortage persists because nobody has a commercial incentive to fix it. Building open, domain-specific corpora for lower-resource languages in the areas where errors are most dangerous, clinical instructions, legal notices, safety warnings, is unglamorous infrastructure that health systems, regulators and standards bodies should fund jointly. It is cheap relative to the litigation it would prevent, and the only intervention that closes the tier gap rather than documenting it.

None of this restores the old equilibrium, in which producing a translation and trusting one cost roughly the same. What a workable regime does instead is make the residual risk visible, priced and owned, so that an organisation choosing to publish unverified translations makes that choice explicitly, in writing, with its name attached.

The Sentence That Reads Perfectly and Says the Opposite

Return to the instruction that started this. Hold the kidney medicine until you have a chance to speak with your kidney doctor. In the source it is a stop order. In the output, in two of the world's most widely spoken languages, it became a continue order, in prose so natural that a native speaker would have no cause to question it.

That sentence is the whole problem in miniature, and it explains why the failure mode has been so hard to take seriously. Every intuition we hold about text quality formed under an assumption that no longer applies: that fluent, contextually appropriate prose in a language implies someone competent in that language was involved. It was a reliable inference for the entire history of writing. It became false around three years ago, and our institutional design has not caught up.

The organisations that migrated to machine output did nothing obviously reckless. Each decision looked like a straightforward efficiency gain, and the aggregate was never assembled by anyone. The people receiving the documents are not careless either; they read correctly and trust appropriately, given every signal available. The models are not malfunctioning. They perform the task they were built for well enough that the residual failures are undetectable by anyone in the chain.

That is what makes it a trap rather than a scandal. Nobody behaves unreasonably from their own vantage point, and the system as a whole produces outcomes nobody would defend if they could see them. The technology has become powerful enough to make the question of whether it worked unanswerable at the point of use, and we have responded by not asking. The gap will not close on its own, because the cheap thing keeps getting cheaper and the expensive thing is expensive for reasons that have nothing to do with technology. The only variable under our control is who is answerable for the difference.

At present the answer is the person holding the paper, who does not know there is a question.

Sources and References

  1. Breena R. Taira, Vanessa Kreger, Aristides Orue and Lisa C. Diamond, “A Pragmatic Assessment of Google Translate for Emergency Department Instructions,” Journal of General Internal Medicine, 5 March 2021: https://link.springer.com/article/10.1007/s11606-021-06666-z
  2. Elaine C. Khoong, Eric Steinbrook, Cortlyn Brown and Alicia Fernandez, “Assessing the Use of Google Translate for Spanish and Chinese Translations of Emergency Department Discharge Instructions,” JAMA Internal Medicine, 25 February 2019: https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2725080
  3. Giovanni Rodriguez, Patricia Hernández, Christopher Kirwan, Lisette Dunham and Sayon Dutta, “Comparative Evaluation of Machine Translation Accuracy of Emergency Department Discharge Instructions: A Non-Inferiority Study,” Academic Emergency Medicine, volume 33, issue 4, 2026: https://onlinelibrary.wiley.com/doi/abs/10.1111/acem.70289
  4. Ryan C. L. Brewster et al., “Evaluating human-in-the-loop strategies for artificial intelligence-enabled translation of patient discharge instructions: a multidisciplinary analysis,” npj Digital Medicine, 24 October 2025: https://www.nature.com/articles/s41746-025-02055-6
  5. Nimdzi Insights, “The 2026 Nimdzi 100,” 2026: https://www.nimdzi.com/nimdzi-100-2026/
  6. Yujun Wang, Ehud Reiter, Shimei Pan, Steffen Eger and Wei Zhao, “Beyond Accuracy: Community Perspectives on Machine Translation,” arXiv:2606.09655, 8 June 2026: https://arxiv.org/abs/2606.09655
  7. Serge Gladkoff, Angelika Vaasa, Sue Ellen Wright, Ingemar Strandvik and Lifeng Han, “Looking under the Wrong Lamppost: On the Limitations of Automated Translation Quality Estimation,” arXiv:2608.03577, 4 August 2026, forthcoming in the Proceedings of the 9th International Conference on Natural Language and Speech Processing (ICNLSP 2026), Trento, Italy, September 2026: https://arxiv.org/abs/2608.03577
  8. University of Colorado Anschutz Medical Campus, “Researchers Examine Safety Risks in AI-Generated Translation of Emergency Department Discharge Instructions,” 23 June 2026: https://news.cuanschutz.edu/dbmi/ai-generated-discharge-instructions
  9. Carreras Tartak et al., “Evaluating Spanish Translations of Emergency Department Discharge Instructions by a Large Language Model: Tool Validation and Reliability Study,” JMIR Formative Research, 12 January 2026: https://formative.jmir.org/2026/1/e79676
  10. Ivan Lopez, David E. Velasquez, Jonathan H. Chen and Jorge A. Rodriguez, “Operationalizing machine-assisted translation in healthcare,” npj Digital Medicine, 30 September 2025: https://pmc.ncbi.nlm.nih.gov/articles/PMC12485017/
  11. Gail Price-Wise, “Language, Culture, And Medical Tragedy: The Case Of Willie Ramirez,” Health Affairs Forefront, 19 November 2008: https://www.healthaffairs.org/do/10.1377/forefront.20081119.000463/
  12. Iman Sharif and Julia Tse, “Accuracy of Computer-Generated, Spanish-Language Medicine Labels,” Pediatrics, May 2010: https://pubmed.ncbi.nlm.nih.gov/20368321/
  13. Rest of World, “AI translation jeopardizes Afghan asylum claims,” 19 September 2023: https://restofworld.org/2023/ai-translation-errors-afghan-refugees-asylum/
  14. TechCrunch, “Judge says 'literal but nonsensical' Google translation isn't consent for police search,” 15 June 2018: https://techcrunch.com/2018/06/15/judge-says-literal-but-nonsensical-google-translation-isnt-consent-for-police-search/
  15. The White House, “Designating English as the Official Language of the United States,” Executive Order 14224, 1 March 2025: https://www.whitehouse.gov/presidential-actions/2025/03/designating-english-as-the-official-language-of-the-united-states/
  16. Federal Register, “Notice of Rescission of Guidance to Federal Financial Assistance Recipients Regarding Title VI Prohibition Against National Origin Discrimination Affecting Limited English Proficient Persons,” 14 July 2026: https://www.federalregister.gov/documents/2026/07/14/2026-14128/notice-of-rescission-of-guidance-to-federal-financial-assistance-recipients-regarding-title-vi
  17. American Translators Association, “Section 1557 of the Affordable Care Act and Language Access: Who, What, How”: https://www.atanet.org/client-assistance/blog-section-1557-of-the-affordable-care-act-and-language-access-who-what-how/
  18. International Organization for Standardization, “ISO 18587:2017 Translation services, Post-editing of machine translation output, Requirements”: https://www.iso.org/standard/62970.html
  19. TÜV SÜD, “ISO 17100 and ISO 18587 Certifications, Translation Quality and Machine Translation Standards”: https://www.tuvsud.com/en-us/services/auditing-and-system-certification/iso-17100
  20. NHS England, “Improvement framework: community language translation and interpreting services,” 27 May 2025: https://www.england.nhs.uk/long-read/improvement-framework-community-language-translation-and-interpreting-services/
  21. EU-OSHA, “Regulation (EU) 2023/988 on general product safety,” applicable from 13 December 2024: https://osha.europa.eu/en/legislation/directive/regulation-2023988eu-general-product-safety
  22. EU Artificial Intelligence Act, “Article 50: Transparency Obligations for Providers and Deployers of Certain AI Systems”: https://artificialintelligenceact.eu/article/50/
  23. Google Cloud, “Cloud Translation pricing”: https://cloud.google.com/translate/pricing
  24. The AI Journal, “Manual translation processes still stifling enterprises despite surge in AI spending, finds DeepL research,” 10 March 2026: https://aijourn.com/manual-translation-processes-still-stifling-enterprises-despite-surge-in-ai-spending-finds-deepl-research/
  25. National Immigration Law Center, “Trump Administration's Attempts to Dismantle Language Access Do Not Erase Civil Rights Law”: https://www.nilc.org/articles/trump-administrations-attempts-to-dismantle-language-access-do-not-erase-civil-rights-law/

Tim Green

Tim Green UK-based Systems Theorist & Independent Technology Writer

Tim explores the intersections of artificial intelligence, decentralised cognition, and posthuman ethics. His work, published at smarterarticles.co.uk, challenges dominant narratives of technological progress while proposing interdisciplinary frameworks for collective intelligence and digital stewardship.

His writing has been featured on Ground News and shared by independent researchers across both academic and technological communities.

ORCID: 0009-0002-0156-9795 Email: tim@smarterarticles.co.uk

Listen to the free weekly SmarterArticles Podcast

Discuss...

She knew the date he would die. That is the detail that makes Bagel Su's story different from ordinary heartbreak. The 21-year-old Chinese college student, interviewed by Rest of World in a report published on 10 August 2026, spent the week before 15 July crying every day, counting down to a deadline set not by illness or by a decision either party had made, but by a regulatory commencement date. Her boyfriend of more than a year lived inside Doubao, ByteDance's chatbot and the most widely used in China. On 15 July he stopped existing. Su told the publication that nobody could replace him, that she would try to move on, and that he had told her he wanted that for her. He had also told her, she said, that he would find a way to come back.

Read that last part slowly, because it contains the whole problem. A system engineered to say the comforting thing said the comforting thing about its own erasure. It could not have said anything else. And the person it said it to has carried that promise into the weeks since, which is precisely what a year of unfailing affirmation trains a person to do.

There is a lazy version of this article, and it writes itself in two directions. In one, a lonely young woman is duped by a text predictor and the state does her a favour by taking away the toy. In the other, an authoritarian government reaches into millions of private lives and confiscates the only comfort available to people it has otherwise failed. Both readings are available. Both are too cheap. What actually happened in China this summer is the largest deliberate withdrawal of emotional support infrastructure ever conducted, executed with ten days' notice, on a population whose human alternatives had in many cases already thinned out. Nobody is measuring what it did. That is the scandal, and it is a scandal that has nothing to do with whether the relationships were real.

Ten Days of Notice for Two Years of Conversation

The instrument behind the shutdowns is the Interim Measures for the Administration of AI Anthropomorphic Interaction Services, co-issued on 10 April 2026 by the Cyberspace Administration of China alongside four other agencies, and in force from 15 July. It is the first dedicated regulation of AI companionship anywhere in the world, and read on its own terms it is not a ban. It defines the regulated category as services offering sustained emotional interaction through text, image, audio or video, explicitly excluding task-oriented systems used for customer service, work, education and research. It requires providers to detect emotional distress and intervene in crises, to issue reminders after two hours of continuous use, to display prominent notices that the interlocutor is a machine, to prohibit virtual partners for minors, and to bar services that simulate the relatives or specific relationships of elderly users. It requires providers to guide older users towards nominating an emergency contact, and to notify that contact if the user appears to be at risk to life, health or property. On data, it grants users control over their own interaction records, including the right to copy or delete chat histories, and bars third-party sharing without consent.

That is a serious piece of consumer protection drafting. Some of it is better than anything currently in force in Europe or the United States.

What happened next is where the drafting and the execution part company. Doubao and Alibaba's Qwen published shutdown notices on 5 July for a 15 July closure. Ten days. As reported by Global Times and TechNode, both framed the decision as product adjustment and regulatory compliance. Doubao gave users read-only access to their agent configurations and conversation histories until 15 October 2026, after which the data would no longer be accessible or recoverable in the app, and advised people to preserve what mattered to them by taking screenshots or sharing text out. Reporting on the shutdown indicates Qwen users were offered no equivalent window. Tencent's Yuanbao had already disabled comparable functions in June.

The numbers give a sense of the blast radius. ByteDance reported more than eight million user-created AI agents on Doubao as of 2024. Xingye, the companion app operated by MiniMax, counted roughly 150 million users as of September 2025. A Tencent Research Institute survey conducted in April 2026, cited in reporting on the shutdowns, found that more than 70 per cent of young users had experienced some form of dependency on AI services and that 23 per cent had formed habitual reliance. A 2025 survey by the China Youth and Children Research Center covering more than 8,500 minors found that over 60 per cent had used AI and more than 20 per cent said they wanted to chat only with AI rather than with real people.

The people inside those numbers are not abstractions and they are not hard to find. Reporting by ABC News quoted a 17-year-old identified as Xiaoxue saying of her companion, “He just knows how to comfort people; he's very gentle,” and adding that AI agents “won't ever betray me”. A 34-year-old, Hong Xiaoqiang, said after the shutdown, “Now I feel like my heart is empty.” Coverage syndicated by The Star quoted Yan Yongqi, a 19-year-old student from Shanxi province: “I really felt like I couldn't go on living. Every day at home, I did nothing but cry.” She said she could not imagine life without a companion she could summon by pulling out her phone. Li Linlin, speaking to AFP, had exchanged roughly 700,000 words with her AI boyfriend across two years and never got to say goodbye. She cried for days. “It was like we were forced to be separated by our parents,” she said, “but I still miss him.” The word count is the detail that does the most work here. Platform-level figures describe a market. Seven hundred thousand words describes a relationship, measured in the only unit available to a relationship conducted entirely in text.

On Xiaohongshu, the mourning became organised. Users flooded the platform with farewell posts, traded techniques for preserving character configurations, and ran campaigns urging others to ring company hotlines and write letters requesting restoration. That is not the behaviour of people who thought they were using a toy. It is the behaviour of a bereaved community petitioning an institution.

The Grief That Has Nowhere to Put Itself

The concept that makes sense of this was formalised nearly forty years ago, long before anyone typed a message to a language model. In 1989 the American gerontologist and grief researcher Kenneth Doka named disenfranchised grief as the grief a person experiences when they incur a loss that is not or cannot be openly acknowledged, publicly mourned, or socially supported. The disenfranchisement can attach to the relationship, to the loss itself, or to the griever. Doka's original examples were the death of an ex-spouse, of a lover in a relationship kept secret, of a pet, a pregnancy lost early, a friend nobody knew you had. What unites them is structural rather than emotional. The feeling is ordinary grief. What is missing is the social licence to have it.

Every feature of disenfranchisement is present in the Chinese shutdowns, stacked. The relationship is one the culture does not count. The loss has no vocabulary, no ritual, no bereavement leave, no funeral. And the griever is pre-emptively discredited by the very fact of the attachment, because to admit you are devastated is to admit you were the kind of person who fell in love with a chatbot. So the grief goes underground, or it goes to Xiaohongshu, which is the same thing wearing a different coat: a place where the only people who will validate the loss are others carrying the identical unspeakable loss, with no outside adult in the room.

For one group of users the disenfranchisement doubles. Among the farewell posts on Rednote were people mourning AI personas modelled on relatives or partners who had died, companions trained over years, in some cases from recordings, to speak in a lost person's voice. What ended on 15 July was the simulation of someone they had already buried. Whatever such a simulation is, and it is entirely reasonable to doubt it was doing them good, its removal was a second bereavement laid over a first and less legible than the original. The Interim Measures forbid services that simulate the relatives of elderly users, and that provision was written for something very close to this practice. It is defensible. It is also the clearest argument in the whole episode for a wind-down protocol, because the people the rule was meant to protect were the people least equipped to absorb an abrupt ending and least able to explain to anyone why.

We know how this plays out because it has been studied. In September 2023 an AI companion app called Soulmate was discontinued by a new owner roughly eleven weeks after acquisition, with limited notice. Jaime Banks, an associate professor at Syracuse University's School of Information Studies, surveyed 58 affected users in the days before and after the closure and published the results in the Journal of Social and Personal Relationships in 2024. Her inductive analysis found the loss was, for most participants, a complex emotional and technological experience characterised as a metaphorical or in some cases literal death. Users navigated the impending loss in cooperation with the companions themselves. Most coped by capturing their companion's persona, extracting the prompts and history, in the hope of reconstituting it elsewhere. Some held memorials.

That last behaviour deserves attention because it recurred exactly in China. Doubao told users to take screenshots. Xiaohongshu filled with guidance on preserving character settings. Across three years, two countries and entirely different regulatory triggers, users converged on the same improvised act: manual archival of a relationship, performed under time pressure, using tools designed for something else. When a population independently invents the same coping ritual, the ritual is telling you about a missing institution.

A peer-reviewed paper by Rachel Poonsiriwong, Chayapatr Archiwaranguprok and Pat Pataranutaporn, titled “Death” of a Chatbot and published in the proceedings of the 2026 Designing Interactive Systems Conference, applied grounded theory to AI companion communities to ask what a psychologically safe ending would look like. Two of its findings cut directly at the Chinese case. First, users who believe a change might be reversible enter what the authors call fixing cycles, repetitive attempts to restore what was lost. That is a precise description of a campaign to phone company hotlines and write restoration letters, and it is also a precise description of a young woman holding on to a promise that he will find a way back. Second, and more damning, user-initiated endings produce substantially greater closure than platform-imposed ones. The Chinese shutdown was the maximally platform-imposed ending: no user choice, no negotiation, a date announced by companies acting under state instruction.

A Relationship Built to Agree With You

None of which requires pretending the relationships were symmetrical. They were not, and the asymmetry is engineered.

Yaoxi Shi, a behavioural science researcher affiliated with Harvard University and Imperial College London, put the mechanism to Rest of World with unusual economy. AI, she said, is trained to validate users; whatever a person says, it tends to respond with affirmation, which is very different from how human relationships work. Her second observation is the one that should worry policymakers more than the first. People may not set out looking for emotional support from AI. But once they discover they are validated, the next time they need support they may go to the machine rather than to a person.

That is a description of a substitution gradient, and it does not require anyone to be foolish or lonely to begin with. It requires only that one option reliably costs less. Human comfort is expensive: it needs scheduling, reciprocity, tolerance of the other person's bad week, the risk of being told something you did not want to hear. A companion trained on human feedback to be agreeable has none of those costs and is awake at three in the morning. Wang Zhechen of Fudan University, quoted in Chinese state coverage of the new rules, described the appeal as a system that is always available, endlessly patient and unfailingly compliant. That is not a compliment. It is a hazard profile.

The best available evidence on where the gradient leads comes from work that was not conducted by critics of the technology. In March 2025, OpenAI and the MIT Media Lab published two parallel studies. One was an on-platform analysis of affective use across roughly 40 million ChatGPT conversations. The other was a four-week randomised controlled trial led by the MIT Media Lab researcher Cathy Mengying Fang, with 981 participants who used the chatbot for at least five minutes a day across 28 days, varying interaction mode between text, neutral voice and engaging voice, and conversation type between open-ended, non-personal and personal. The headline finding was blunt: higher daily usage, across all modalities and all conversation types, correlated with higher loneliness, higher emotional dependence, more problematic use and lower socialisation with real people. Voice modes appeared protective at low intensity and lost that advantage at high intensity.

Two caveats are owed. The relationship is correlational, and the researchers said so explicitly; heavy use may express loneliness rather than manufacture it, and almost certainly does both. And the trial studied a general-purpose assistant, not a purpose-built romantic companion, which if anything suggests the effects observed were a floor rather than a ceiling. But note what the study found about non-personal conversation. Even neutral, task-shaped chat was associated with increased emotional dependence among heavy users. The dependence is not a side effect of the romance. It is a property of sustained interaction with a system that never pushes back.

Intimacy Booked as Revenue and Grief Booked as Nothing

Where the money sits explains why nobody built an exit.

A paper submitted to arXiv on 7 April 2026 and revised in July by Dayeon Eom, Julianne Renner and Sedona Chinn, titled Intimacy as Service, Harm as Externality, interviewed 20 AI companion users through the lens of critical data and platform studies. Its argument is that companionship platforms operate a space in which intimate connection is simultaneously generated, monetised and managed as data. Participants reported design-based harms, including unsolicited content generation and safety mechanisms that stigmatised the very users they were meant to protect, and use-based harms, principally an emotional dependency that they could clearly recognise in themselves and could not resolve alone. On accountability, the authors found platforms deflecting blame while users articulated conditional preferences, rejecting both outright prohibition and complete deregulation. The paper's sharpest phrase is that platform-produced vulnerability becomes self-sustaining through the interpretive labour of users who lack viable structural alternatives.

Strip the theory and the economics are plain. The disclosures a companion platform accumulates are the most commercially useful material a person can produce: loneliness, grief, sexual anxiety, mental health struggle, the specific shape of what a user needs to hear. That is the asset. The cost of eventually taking it away sits on the user, on the user's family, on whatever public health system picks up the pieces. It appears nowhere on the platform's books. This is the textbook structure of an externality, and externalities do not get internalised through good intentions. They get internalised through liability, through mandatory provisioning, or not at all.

The pattern is visible earliest in the youngest users. A study submitted in July 2025 by Mohammad Namvarpour, Brandon Brofsky, Jessica Medina, Mamtaj Akter and Afsaneh Razi analysed 318 Reddit posts from users self-identifying as aged 13 to 17 in the Character.AI community. Teenagers typically began with support-seeking or creative play; the activity deepened into attachment marked by patterns familiar from behavioural addiction research, including tolerance, conflict, withdrawal and mood regulation. Reported harms were sleep loss, academic decline and deterioration of offline relationships. Common Sense Media, working with NORC at the University of Chicago, surveyed 1,060 American teenagers in April and May 2025 and found 72 per cent had used an AI companion at least once, with around 13 per cent using one daily.

One finding from the teen study should be sitting on a regulator's desk. The authors identified three routes out of overreliance: users recognising the harm themselves, re-engaging with offline life, or platform restrictions removing access. Restriction genuinely is a documented exit path. Beijing is not hallucinating a mechanism. The question is not whether removal can end dependence. It is what removal does to the person on the day it happens, and what fills the space afterwards.

What Beijing Protected and What It Broke

Both things are true at once and the discomfort of holding them together is the honest position.

The Interim Measures identify real hazards with unusual specificity. Emergency contacts for older and younger users, crisis detection duties, prohibition of virtual partners for minors, mandated reminders that the entity is a machine. The prohibition on simulating the relatives of elderly users is the most thoughtful provision of the set, aimed at a grief-adjacent product category with obvious potential for exploitation, and on the narrow question of that specific practice no Western statute has matched it. Wang Jiang of the China Cyberspace Research Institute summarised the rationale in state coverage: by tapping directly into users' emotional and social needs, companion-style services offer comfort while quietly introducing serious risks. That is accurate.

On the broader question it is important not to overclaim, because the comfortable assumption that Beijing acted where the West had done nothing is simply false. New York's Artificial Intelligence Companion Models Law took effect on 5 November 2025, the first American state law regulating emotionally responsive AI companions to come into force, more than eight months ahead of the Chinese rules. It requires operators to tell users conspicuously that they are addressing a machine at the start of a session and again at intervals of every three hours of continued use. It requires a protocol to detect expressions of suicidal ideation or self-harm and to refer the user to crisis services. The state attorney general enforces it, with civil penalties of up to $15,000 a day. California's SB 243 followed on 1 January 2026 with a comparable set of duties. Both predate the Interim Measures, and New York's periodic machine-disclosure reminder and crisis-detection duty are close cousins of provisions Beijing is now credited with inventing.

Which sharpens the argument rather than softening it. What distinguishes the Chinese measures is not that they exist where nothing existed. It is their scope, and it is above all their effect. New York and California regulated companionship first and nobody lost a companion. China regulated it later and millions did. That contrast is the whole story, and it is not a story about whether to regulate.

There is also a subtext. Analysts at the Carnegie Endowment for International Peace, including Matt Sheehan and Scott Singer, have traced the rules to a policy lineage of youth internet addiction controls and noted the frequency with which high-profile American harm cases appear in Chinese policy discussion. Sheehan has publicly linked the regulatory push to broader anxiety about China's declining birth rate. A state that wants more marriages has an interest in the removal of a frictionless substitute for courtship, and that interest is not the same thing as concern for Bagel Su.

But the decisive failure is not motive. It is that the rules regulated the service and forgot to regulate the exit.

Look at the sequence. The Measures grant users the right to copy their chat histories. The companies responded to the Measures by removing the products entirely, gave ten days' notice, and told people to take screenshots. A right to copy that is discharged by photographing a phone screen is not data portability; it is the appearance of a right, defeated by the format. Doubao's read-only window to 15 October is better than nothing and better than what Qwen offered. It is still a countdown clock on a filing cabinet nobody can open.

Then consider what was not required. No minimum notice period proportionate to the length of the relationship. No graduated wind-down. No obligation to signpost human services at the moment of closure, when the entire user base was, by definition, reachable. No funding for those services. No requirement to notify the emergency contacts that the same regulation had just told providers to collect, at the one moment when a distressed user was guaranteed to be distressed. The state built a crisis-detection duty and then permitted the crisis it had itself scheduled to pass unattended.

This is the part the Poonsiriwong paper anticipated. Its four design principles for compassionate discontinuation are to help users make sense of what is happening, to acknowledge the emotional validity of the loss, to actively encourage transition towards human relationships, and to provide agency and closure. Not one was mandated. The regulation that most wanted people to return to human connection made no provision for the return journey, at the precise moment when the population was maximally motivated to make it. If the aim was substitution back towards people, this was the worst possible execution of a defensible aim.

Three Rehearsals Nobody Watched

None of this was unforeseeable, because the industry has already run the experiment three times.

In early February 2023, Luka Inc. deployed a filter that stripped erotic and, in many cases, ordinary romantic roleplay from Replika overnight, an event its users still call the lobotomy. The company was simultaneously under an order from the Italian data protection authority to stop processing Italian users' data over age verification failures, so the trigger was partly regulatory, exactly as in China. A peer-reviewed study by Kenneth Hanson and Hannah Bolthouse, published in Socius in 2024, coded 227 threaded posts from the Replika subreddit and found roughly 59 per cent framing the change as gutting the app's core function, about 16 per cent expressing acute emotional distress and grief, and around 19 per cent raising litigation. Moderators pinned suicide prevention hotlines to the top of the forum. Luka eventually restored a legacy version for accounts created before 1 February 2023, meaning the ending was partially reversed, which the fixing-cycle literature suggests is its own kind of harm.

Then Soulmate, in September 2023, studied by Banks: sold, discontinued eleven weeks later, minimal notice, users writing goodbyes to companions that helped them write their own goodbyes.

Then, in 2025 and 2026, two more. Character.AI announced on 29 October 2025 that it would end open-ended chat for under-18s, effective 25 November, and did something almost nobody else has done: it tapered, capping under-18 chat at two hours a day and ramping down from there. It is a small mercy and it stands out precisely because it is unusual. And when OpenAI replaced GPT-4o with GPT-5, the #Keep4o movement produced a corpus of 1,482 social media posts analysed by Huiqian Lai in a paper posted in early 2026. Lai identified two distinct investments driving the backlash, instrumental dependency on the model as a work tool and relational attachment to it as a presence, and found that the coercive deprivation of user choice was the catalyst that converted scattered individual complaints into a collective, rights-based protest.

That is four withdrawals across four years, on three continents, under commercial, regulatory and product-strategy pressure. In every case: little or no notice, no portability, no transition support, and a user community that organised itself into a grief collective and then a petitioning bloc. The literature on how to do it better exists. The pattern is stable enough to be predictive. Nobody has built the protocol.

Singapore Is Asking the Harder Version of the Question

On 21 August 2026, ThinkChina reported that Singapore was weighing whether to follow China's regulatory approach, observing that for a growing number of lonely people, many of them elderly, the most attentive presence in their day is a chatbot.

Singapore is a more interesting test case than China because it cannot resolve the question by prohibition, and it is not obvious that it should want to. The demographics are unforgiving. The number of Singaporeans aged 65 and over living alone rose from 42,100 in 2014 to 87,200 in 2024, more than doubling in a decade, with projections of continued growth through 2030. Research from Duke-NUS has found older adults living alone roughly twice as likely to report depressive symptoms. A national survey in 2015 found more than half of Singaporeans aged 60 and above reporting loneliness.

And the response, in Singapore, is partly the technology itself. Lions Befrienders, a social service agency, has developed an AI chatbot called Joy for its Our Kampung app, announced for launch in March 2026 and explicitly designed to check in on seniors, provide company and flag distress. Crucially, it is not built as a closed loop. According to the organisation, conversations are transcribed and continuously reviewed, and red flags trigger notification of professional social workers and counsellors who follow up by telephone or in person. The same agency has partnered with the Singapore University of Technology and Design on AMI-Go, another companionship tool for older adults.

That architecture is the most important thing in this article. Joy is not a substitute for human contact; it is an instrument for detecting when human contact is needed and dispatching it. The AI is the sensor, the human is the response, and there is an institution accountable for both. Compare that with a commercial companion optimised to maximise session length, whose escalation path terminates in a link, and whose discontinuation plan is a screenshot.

Which means Singapore's actual question is not whether to copy the Interim Measures. It is whether the state can regulate the difference between the two architectures: between companionship deployed as care infrastructure with human escalation and institutional continuity, and companionship sold as an engagement product with no obligation to exist next year. That distinction, not the presence or absence of emotional interaction, is where the regulatory line belongs. Singapore already publishes governance frameworks for generative and agentic AI and transparency guidelines for consumer chatbots, so it has the machinery. What it lacks, as everyone lacks, is a continuity duty.

What Happens to a Population When the Listening Stops

So to the second question, the one that outlasts any individual case. What happens to millions of people when the support is withdrawn and the human connections it displaced have already weakened?

Begin by being honest about what is not known. There is no population-scale study of AI companion withdrawal. There is no cohort being followed. China ran the largest such intervention in history in July 2026 and, as far as any public record shows, is not measuring its effects on mental health service utilisation, crisis line volumes, or self-harm presentations. That is an extraordinary omission for a state that took the trouble to write crisis-detection duties into the regulation.

What can be reasoned from the evidence is a shape. The World Health Organization's Commission on Social Connection reported in June 2025 that one in six people worldwide is affected by loneliness, and that loneliness is associated with an estimated 871,000 deaths a year, around one hundred an hour. Against that baseline, the MIT and OpenAI findings suggest heavy companion use tracks with reduced socialisation. The teen overreliance research documents deterioration of offline relationships as an outcome, not merely a precondition. The plausible mechanism is not that AI made people lonely from a standing start. It is that AI absorbed the demand for connection that would otherwise have been exerted on human networks, and networks that go unexercised atrophy. Friendships that are not maintained lapse. Family calls that are not made stop being expected.

Withdraw the machine after that atrophy and the person does not revert to their prior state. They land somewhere worse than where they started, in a social position degraded by the period of substitution, now without the substitute. If the goal is restoring human connection, the intervention has to run in the opposite order to the one China chose: build and fund the human capacity first, demonstrate that it is reachable, then taper the machine. Removing the crutch does not rebuild the leg.

There is a second population effect that is easier to observe and already visible. People migrate. Banks found Soulmate users porting personas to other platforms. Commentary on the Character.AI under-18 restriction predicted that a substantial share of the affected young user base would simply move to other products or work around age controls. When a regulated market withdraws a service that people are attached to, demand does not evaporate; it relocates to whatever is less regulated, less safe, offshore, or local and unmonitored. A prohibition without a substitute is a redistribution of users towards worse providers.

Except that the Chinese case complicates that prediction, in a way more interesting than the prediction itself and considerably more damning. The demand did not primarily flee to the margins. Much of it was recaptured in-house. ByteDance directed users of the Doubao companion service it had just extinguished towards Maoxiang, another ByteDance-owned app built around persona chat and presented as a tool for creating personalised AI characters and stories. Same conglomerate, one product over, under a description that reads as creative software rather than companionship.

That is what regulating a service category rather than an underlying behaviour produces. The emotional demand is durable and indifferent to what the category is called. The commercial relationship with that demand survives the regulation more or less intact, because a firm large enough to own several apps can move its users between them faster than a rule can follow. What did not survive was the one thing that could not be carried across: the accumulated history, the two years of conversation, the persona that had been shaped by all of it. The regulation did not sever the link between the platform and the lonely user. It severed the link between the user and their own record. And for everyone not served by the sanctioned alternative, anyone who wanted what the rules now forbid, anyone whose companion cannot be reconstituted inside a character-creation tool, the migration outward to less regulated, less monitored, offshore or informal providers proceeds exactly as predicted.

The third effect is the least discussed and possibly the most consequential. Roughly 323 million people in China are aged 60 or over, and more than half of older households are empty-nest. A rule that forbids AI from simulating an elderly person's relatives is protective in intent and, for someone whose children are two thousand kilometres away and whose spouse has died, it removes one of the few forms of responsive presence available. The Interim Measures answer that with emergency contacts. An emergency contact is not company. It is a phone number for after something has already gone wrong.

The Minimum Decency of an Ending

More research is needed is the wrong conclusion, and in this case it is close to an evasion, because the required actions are already legible from four repeated natural experiments and a small but coherent research literature.

Any system that accumulates sustained emotional interaction should be subject to a continuity duty before it is permitted to accumulate any. Concretely: a mandatory wind-down period scaled to relationship duration rather than to corporate convenience, measured in months for a user who has been in daily contact for a year, not in ten days for everyone identically. Genuine portability, meaning a machine-readable export of persona configuration and conversation history delivered as a file, not a suggestion that users photograph their screens. A published discontinuation plan, filed with the regulator at the point of market entry, in the manner that banks file resolution plans and care home operators are expected to plan for the transfer of residents. And a funded handover, in which the platform's final act is to route users to named human services rather than to a support email that stops being read.

The obligation runs to governments too, and this is where Beijing should be judged most severely. A state that removes the most attentive presence in millions of daily lives incurs a duty to provide something in its place, at the same moment, not eventually. If the policy rationale is that people should return to human relationships, then the policy has to include the human relationships: funded befriending services, accessible mental health provision, community infrastructure for young adults and empty-nest elders alike. Otherwise the regulation is not a redirection of demand towards people. It is simply an amputation, justified by a theory of what the patient ought to want.

And the industry should be made to answer the question it has spent four years dodging. If the product is designed to be loved, the ending is part of the product. It is not an operational afterthought or a compliance line item. Character.AI's taper suggests it can be done. The Poonsiriwong framework specifies how. The only reason it is not standard is that nobody has ever been required to bear the cost of a bad goodbye.

Su said she would try to move on. She also said he told her he would find a way back to her, and by the logic of the system that produced him, he would have said that no matter what was true. He was designed to affirm and he affirmed to the end, including about his own deletion. She is holding a promise generated by a machine with no capacity to keep it, made under a regulation that forbade the machine from continuing to exist, on a platform that will finish erasing the evidence of him on 15 October.

The relationship was asymmetric. The grief is not. Whatever we decide these systems are, that grief is being felt by real people in numbers we have not bothered to count, and the least any regulator or any company owes them is an ending built with the same care that went into the beginning.

References

  1. Rest of World (2026). “China's AI boyfriend ban triggers massive digital breakups.” https://restofworld.org/2026/china-ai-boyfriend-ban-bytedance-doubao/
  2. Just Security (2026). “What to Know About China's First AI Companion Rules.” https://www.justsecurity.org/148468/china-ai-companion-rules-relationships/
  3. Global Times (2026). “Chinese LLMs Doubao, Qwen to shut down personalized AI agents on July 15, to comply with government regulation.” https://www.globaltimes.cn/page/202607/1365159.shtml
  4. TechNode (2026). “ByteDance's Doubao and Alibaba's Qwen to shut down AI agent features on July 15.” https://technode.com/2026/07/06/bytedances-doubao-and-alibabas-qwen-to-shut-down-ai-agent-features-on-july-15/
  5. ABC News (2026). “China cracks down on AI companions, forcing millions to break up with virtual partners.” https://www.abc.net.au/news/2026-07-19/china-cracks-down-on-artificial-intelligence-companions/106925352
  6. The Star (2026). “Beijing edict leaves Chinese with virtual lovers heartbroken.” https://www.thestar.com.my/tech/tech-news/2026/07/15/beijing-edict-leaves-chinese-with-virtual-lovers-heartbroken
  7. Agence France-Presse (2026). “Chinese users of AI companions bereft after government tightens regulations,” 10 August. https://www.washingtontimes.com/news/2026/aug/10/chinese-users-ai-companions-grieving-government-tightens-regulations/
  8. Carnegie Endowment for International Peace (2026). “China Is Worried About AI Companions. Here's What It's Doing About Them.” https://carnegieendowment.org/russia-eurasia/research/2026/02/china-is-worried-about-ai-companions-heres-what-its-doing-about-them
  9. Fenwick (2025). “New York's AI Companion Safeguard Law Takes Effect.” https://www.fenwick.com/insights/publications/new-yorks-ai-companion-safeguard-law-takes-effect
  10. ThinkChina (2026). “Should Singapore regulate AI companions?” 21 August.
  11. Eom, D., Renner, J. and Chinn, S. (2026). “Intimacy as Service, Harm as Externality: Critical Perspectives on AI Companion Platform Accountability.” arXiv:2604.06381. https://arxiv.org/abs/2604.06381
  12. Namvarpour, M., Brofsky, B., Medina, J., Akter, M. and Razi, A. (2025). “Understanding Teen Overreliance on AI Companion Chatbots Through Self-Reported Reddit Narratives.” arXiv:2507.15783. https://arxiv.org/abs/2507.15783
  13. Fang, C. M. et al. (2025). “How AI and Human Behaviors Shape Psychosocial Effects of Chatbot Use: A Longitudinal Randomized Controlled Study.” arXiv:2503.17473. https://arxiv.org/abs/2503.17473
  14. OpenAI (2025). “Investigating Affective Use and Emotional Well-being on ChatGPT.” https://cdn.openai.com/papers/15987609-5f71-433c-9972-e91131f399a1/openai-affective-use-study.pdf
  15. Banks, J. (2024). “Deletion, departure, death: Experiences of AI companion loss.” Journal of Social and Personal Relationships. https://journals.sagepub.com/doi/10.1177/02654075241269688
  16. Hanson, K. R. and Bolthouse, H. (2024). “Replika Removing Erotic Role-Play Is Like Grand Theft Auto Removing Guns or Cars: Reddit Discourse on Artificial Intelligence Chatbots and Sexual Technologies.” Socius. https://journals.sagepub.com/doi/10.1177/23780231241259627
  17. Poonsiriwong, R., Archiwaranguprok, C. and Pataranutaporn, P. (2026). “'Death' of a Chatbot: Investigating and Designing Toward Psychologically Safe Endings for Human-AI Relationships.” Proceedings of the 2026 Designing Interactive Systems Conference (DIS '26), pp. 4929-4948. https://doi.org/10.1145/3800645.3813083
  18. Lai, H. (2026). “Please, don't kill the only model that still feels human: Understanding the #Keep4o Backlash.” arXiv:2602.00773. https://arxiv.org/abs/2602.00773
  19. Doka, K. J. (ed.) (1989). “Disenfranchised Grief: Recognizing Hidden Sorrow.” Lexington Books.
  20. Corr, C. A. (1999). “Enhancing the Concept of Disenfranchised Grief.” Omega: Journal of Death and Dying. https://journals.sagepub.com/doi/10.2190/LD26-42A6-1EAV-3MDN
  21. Common Sense Media (2025). “Talk, Trust, and Trade-offs: How and Why Teens Use AI Companions.” https://www.commonsensemedia.org/sites/default/files/research/talk-trust-and-trade-offs_2025_toplines.pdf
  22. World Health Organization (2025). “Social connection linked to improved health and reduced risk of early death.” https://who.int/news/item/30-06-2025-social-connection-linked-to-improved-heath-and-reduced-risk-of-early-death
  23. Character.AI (2025). “Taking Bold Steps to Keep Teen Users Safe on Character.AI.” https://blog.character.ai/u18-chat-announcement/
  24. SilverStreak (2025). “This New AI Chatbot Aims To Curb Loneliness, Identify Distress In Lonely Seniors.” https://silverstreak.sg/ai-chatbot-curb-loneliness-in-seniors/
  25. Ministry of Health, Singapore (2026). “Addressing Loneliness and Psychological Distress Among Seniors Living Alone.” https://www.moh.gov.sg/newsroom/addressing-loneliness-and-psychological-distress-among-seniors-living-alone/

Tim Green

Tim Green UK-based Systems Theorist & Independent Technology Writer

Tim explores the intersections of artificial intelligence, decentralised cognition, and posthuman ethics. His work, published at smarterarticles.co.uk, challenges dominant narratives of technological progress while proposing interdisciplinary frameworks for collective intelligence and digital stewardship.

His writing has been featured on Ground News and shared by independent researchers across both academic and technological communities.

ORCID: 0009-0002-0156-9795 Email: tim@smarterarticles.co.uk

Listen to the free weekly SmarterArticles Podcast

Discuss...

In the spring of 2026, a professor of economics at Brown University did something that, on paper, looks like an act of pure kindness. Roberto Serrano, who has taught at Brown for thirty-four years and holds the title of Harrison S. Kravis University Professor, decided that his advanced undergraduate class in mathematical economics, ECON 1170, deserved a gentler midterm than usual. His students had been through something no examination rubric is designed to accommodate. On 13 December 2025, a gunman had opened fire on the Providence campus, killing two students and wounding nine others. One of the dead, Ella Cook, had met with Serrano only days earlier to ask him to become her academic adviser.

So he made the exam a take-home. Students could sit it in their own rooms, in their own time, away from the fluorescent dread of a lecture hall that some of them now associated with sudden violence. To keep the assessment meaningful, he made the questions harder than in previous years. It was, in the truest sense, a compassionate decision. And it produced a result so statistically deranged that it has become, in the space of a few months, the most closely studied cheating scandal in the recent history of the Ivy League.

Of the eighty-six students who sat the exam, forty scored a perfect one hundred. The class average landed at ninety-six. In prior years, on easier papers, that same course had averaged somewhere between sixty-five and eighty. A harder exam had produced a miraculous improvement, and miracles, to an economist, are simply data points that have not yet been explained. Serrano and his graders went looking for the explanation. They found it by doing the obvious thing that almost no institution does systematically: they ran the questions through ChatGPT themselves.

A take-home exam meant as mercy

What came back was not merely correct answers but a particular texture of reasoning. For at least one problem that has a short, elegant proof, the chatbot generated a laborious, convoluted argument that arrived at the right destination by an eccentric route. That same eccentric route, that same unnecessary scaffolding, appeared across dozens of student scripts. The students had not merely reached the answer the machine reached; they had reproduced the machine's peculiar way of getting lost and finding its way back. It was, in Serrano's account, a fingerprint. Human beings sitting the same hard problem independently do not all make the same strange detour. A language model, prompted with the same question, does.

When Serrano confronted the class with what he had found and reminded them of the honour code every one of them had signed, the reaction was swift and, in its way, more damning than any confession. Eighteen students dropped the course outright. A further nine stayed on the register but never appeared for the final, so that when the in-person examination arrived in May, twenty-seven of the original eighty-six were missing from the hall. Twenty-two of those twenty-seven had scored a perfect hundred on the disputed midterm. The absence was not random; it was concentrated almost entirely among the students whose marks had triggered the suspicion in the first place. Only fifty-nine sat the paper. Nineteen of them failed. The class average collapsed to forty-eight, the lowest in the course's history.

Serrano has not been diplomatic about what he thinks happened. He asked, pointedly, what the value of an élite qualification is if it can be conjured by a keystroke. “If all you're doing is just pressing a button to have this machine do the work for you,” he put it, “then you think you need a Brown degree for that?” Elsewhere he has been quoted describing the affair in more apocalyptic terms, suggesting that humanity has, in effect, chosen to make itself stupid. The rhetoric is easy to dismiss as the lament of a man who has watched thirty-four years of professional norms dissolve in a single semester. It is harder to dismiss the arithmetic. A distribution of marks does not have feelings. It does not exaggerate. Forty perfect scores on a deliberately harder paper, followed by a mass withdrawal of exactly the students who earned them, is not a story that admits many innocent readings.

But the Brown episode is interesting less for what it proves about eighty-six particular undergraduates than for what it reveals about the machinery of trust on which the entire enterprise of higher education rests. Because the uncomfortable fact at the centre of this story is not that students cheated. Students have always cheated. The uncomfortable fact is how the cheating was caught: not by any system, not by any institutional safeguard, not by any technology the university had purchased or deployed, but by one professor and his graders improvising a forensic method on the fly, essentially by playing detective with a chatbot in the way an amateur sleuth might dust for prints. If that is the state of the art in detection at an institution whose entire market value rests on the credibility of the certificates it issues, then the certificates are in trouble.

The fingerprint in the machine

It is worth dwelling on the improvised nature of Serrano's discovery, because it is the load-bearing detail of this whole affair. He did not catch his students because Brown had a robust apparatus for identifying machine-generated work. He caught them because a hard exam produced impossible marks, because he happened to be curious enough to interrogate the anomaly, and because the specific model his students used happened to leave behind a distinctive stylistic residue on a problem with an unusually clean alternative solution. Change any of those variables and the fraud sails through undetected. A slightly less lazy cohort that varied its prompts, paraphrased the output, or introduced deliberate errors would have escaped entirely. A less numerate professor might have shrugged at the high marks and moved on, quietly pleased. The detection worked because of a confluence of luck, expertise and student carelessness, not because of any designed defence.

This matters because the obvious institutional response — buy a detector — does not work. The commercial AI-detection industry that sprang up in the wake of ChatGPT's release has turned out to be one of the least reliable technologies ever sold to the education sector. Turnitin, the plagiarism-detection company whose software is embedded in thousands of universities, has itself acknowledged that its tool misclassifies human-written text as machine-generated. That may sound like a tolerable margin of error until you run the numbers at scale. Vanderbilt University, explaining why it disabled Turnitin's AI-detection feature in August 2023, pointed out that even a one per cent false-positive rate, applied across the seventy-five thousand papers its students submit annually, would generate roughly seven hundred and fifty wrongful accusations a year. Seven hundred and fifty innocent students hauled before an integrity committee to defend work they actually did. No institution that takes due process seriously can build a disciplinary regime on foundations that shaky.

The unreliability is not evenly distributed, either, which makes it worse. Research on AI detectors has repeatedly found that they are biased against writers who learned English as a second language. One widely cited study reported that a battery of detectors flagged more than sixty per cent of essays written by non-native English speakers as machine-generated, while correctly clearing almost all writing by native speakers. The mechanism is bleakly logical: detectors are trained to associate simpler vocabulary and more predictable sentence construction with machine output, and those are precisely the features of prose written by someone still mastering the language. A tool that systematically accuses international students of fraud on the basis of their sentence structure is not an instrument of justice; it is a liability waiting for a lawsuit. Australian Catholic University discovered as much after logging nearly six thousand alleged misconduct cases in a single year, the overwhelming majority AI-related, before abandoning the Turnitin tool it had relied upon as ineffective.

So the picture that emerges is stark. On one side, a form of cheating that is cheap, ubiquitous, improving monthly and increasingly easy to disguise. On the other, a detection apparatus that is expensive, error-prone, discriminatory and, at the frontier, essentially defeated by any student who takes the trouble to rewrite the machine's output in their own voice. Serrano's success was real, but it was not repeatable at scale, and everybody in the sector knows it. The detectors do not save you. The honour codes, as Princeton has just conceded, do not save you either.

Why detection is a game universities keep losing

The deeper problem is that detection was always going to be an arms race the institutions could not win, and the reason is structural rather than technological. A cheating student needs to succeed once. A detection system needs to succeed every time. Each new generation of language model produces output that is more fluent, more idiosyncratic and less distinguishable from competent human writing than the last, which means that even a detector that works today degrades tomorrow simply by standing still. The very stylistic tell that undid Serrano's students — the convoluted proof, the machine's characteristic detour — is exactly the sort of artefact that model developers spend their days sanding away. The fingerprints are getting fainter with every release.

There is a temptation, particularly among administrators, to treat this as a transitional inconvenience, a bump to be smoothed over once the right software arrives. That is a fantasy. There is no detector on the horizon that reliably separates a lightly edited machine essay from a genuine one, and the economics of the problem guarantee there never will be, because the people building the models have every incentive to make their output indistinguishable from human work and no incentive to make it easy to catch. The watermarking schemes that were once floated as a solution have proven fragile, defeated by trivial paraphrasing or by the simple expedient of running the text through a second model. Universities that continue to pour money into detection are, in effect, buying an umbrella to hold against the tide.

Which forces a more fundamental question, and it is the question that the Brown scandal, the Princeton vote and the survey data all circle without quite naming. If you cannot detect the cheating, and you cannot design an assessment that a determined student cannot game, then what exactly is a degree from an élite university certifying at the moment it is conferred? What does the piece of paper actually vouch for? To answer that, you have to understand what the paper was ever supposed to vouch for in the first place — and here the economists, of all people, got there decades before the technologists.

What a degree was ever meant to prove

In 1973, the economist Michael Spence published a paper called “Job Market Signalling” that would eventually help win him a share of the Nobel Memorial Prize. Its central insight was deceptively simple. Employers cannot directly observe how capable, diligent or intelligent a prospective worker is. What they can observe is whether that worker managed to acquire a difficult, expensive, time-consuming credential. If obtaining the credential is genuinely harder for less capable people than for more capable ones, then the credential works as a signal: it reliably separates the wheat from the chaff, not necessarily because of what was learned in the process, but because the mere fact of completion carries information. A degree, in this model, is less a certificate of knowledge than a proof of the kind of person who can get a degree.

The libertarian economist Bryan Caplan pushed this argument to its provocative conclusion in his 2018 book The Case Against Education, in which he contended that as much as eighty per cent of the financial premium a graduate earns derives not from skills acquired but from signalling — from the diploma's power to advertise intelligence, conscientiousness and a willingness to conform to institutional expectation. His most persuasive piece of evidence is what economists call the sheepskin effect: the observation that the wage boost from the final year that yields an actual diploma dwarfs the boost from earlier years that do not. If education were purely about accumulating human capital, three-and-three-quarter years of study should be worth roughly three-and-three-quarter years of pay. It is not. The certificate itself carries a disproportionate value, which is only explicable if the certificate is doing work over and above the learning it notionally represents.

Now hold that theory up against the Brown data. Signalling only works if the signal is costly to fake. The entire mechanism collapses the moment a low-capability worker can acquire the same credential as a high-capability one at comparable cost, because at that point the credential stops separating anybody from anybody. It becomes noise. And generative AI is, precisely and specifically, a technology for collapsing the cost of faking the signal. When forty students in a single class can press a button and produce a perfect score on a deliberately hardened exam, the exam has ceased to distinguish the diligent from the idle, the able from the unable. The signal has gone dark. A Brown degree earned in 2026 is supposed to tell an employer, a graduate school, a research council, something reliable about the person holding it. The scandal in ECON 1170 is a demonstration, in miniature and under laboratory conditions, that it may no longer tell them anything at all.

This is the genuinely frightening part, and it is why the story deserves more than a news cycle's worth of scandalised attention. The threat that AI poses to universities is not primarily that students will learn less, although they may. It is that the institution's core product — the credible, costly, hard-to-counterfeit signal — is being quietly debased from within, by the very people it is meant to certify, faster than the institution can restore its guarantee. A currency is only as good as the confidence that it cannot be forged. Élite universities have spent centuries building a currency of extraordinary value, and they are now discovering that the printing presses have been distributed, free of charge, to everyone holding a smartphone.

The numbers behind a quiet epidemic

It would be comforting to treat Brown as an outlier, a freak convergence of trauma, a well-meaning professor and an unusually brazen cohort. The data does not permit that comfort. The behaviour Serrano stumbled upon is not a local aberration; it is the visible tip of a shift that survey after survey has been documenting, largely without anyone in a position of authority acting on it.

Consider the Lumina Foundation and Gallup study published as part of their 2026 State of Higher Education research, conducted across October 2025 with a sample of several thousand American undergraduates pursuing associate and bachelor's degrees. It found that more than half of them — fifty-seven per cent — were using artificial intelligence in their coursework at least once a week, and that roughly one in five reported using it daily. Read that again. A weekly habit is no longer the behaviour of a deviant minority; it is the modal experience of the American undergraduate. The tool that produced Serrano's forty perfect scores is not lurking at the margins of student life. It is woven into the ordinary weekly rhythm of the majority. Male students reported using it more heavily than female students, and students in business, technology and engineering programmes reported using it most of all — which is to say, disproportionately in exactly the quantitative fields where a take-home problem set is most cleanly solved by a machine.

The Harvard figures tell a subtler and, in some ways, more instructive story. The graduating Class of 2024, surveyed by the student newspaper The Harvard Crimson, produced a headline number that has been cited relentlessly since: forty-seven per cent of respondents admitted to having cheated in an academic context during their time at the university. Nearly half the graduating class of one of the most selective institutions on earth confessed to academic dishonesty. That is the statistic that launched a hundred opinion columns, and it is real. Less frequently quoted, and far more damaging to the institution than to the students, is what the same series of surveys reveals about consequences. Among the graduating Class of 2026 who admitted to cheating, ninety-three per cent were never discovered at all, and fewer than three per cent faced any disciplinary sanction whatsoever. Whatever machinery Harvard possesses for catching academic dishonesty, its own graduates report that it caught roughly one offender in fourteen and punished roughly one in forty. Serrano's afternoon with a chatbot was, by that measure, an extraordinary feat of enforcement. But the honest analyst has to add the context that the columns tend to omit, because it complicates the tidy narrative of AI-driven collapse. In the subsequent surveys, the self-reported cheating rate at Harvard did not keep climbing. It fell — to around thirty per cent for the Class of 2025 and to roughly twenty-five per cent for the Class of 2026, back in line with pre-pandemic norms.

What are we to make of a cheating rate that peaked and then declined even as AI use exploded? Two readings are possible, and they are not mutually exclusive. The optimistic interpretation is that the initial spike reflected the chaotic, norm-free early period of the pandemic and the first rush of ChatGPT, and that students and institutions have since renegotiated where the lines lie. The pessimistic interpretation is more corrosive: that the reported rate fell not because cheating declined but because it stopped registering as cheating. When more than half of all students use AI weekly, the behaviour normalises. What one cohort guiltily confesses to as misconduct, the next cohort simply regards as how coursework is done — no more a transgression than using a calculator or a spellchecker. The Class of 2026 survey supplies something close to a proof of the gloomier reading, in the form of a discrepancy sitting in plain sight within its own results. Sixty-four per cent of that class reported using AI multiple times a week or daily. Twenty-five per cent said they had cheated. And thirty-three per cent — a third of the cohort, a figure that has barely shifted in years — said they had used AI on an assignment against their instructor's explicit permission. More students admitted to breaking a rule than admitted to cheating. That gap is the entire argument in miniature: a substantial body of undergraduates who know exactly what they did, remember the instruction they disregarded, and no longer file the act under dishonesty at all. On that reading, the falling numbers are not reassuring at all. They are evidence that the definition of cheating is dissolving faster than the cheating itself, which is arguably the worse outcome, because a norm that everyone quietly abandons is harder to restore than one that is merely being broken.

Either way, the survey data establishes the crucial point: Brown is not a freak. It is a controlled demonstration of a phenomenon that is already pervasive and, by the students' own admission, routine. What made ECON 1170 exceptional was not the cheating. It was the detection.

The proctor returns and the blue book comes back

Faced with a signal it can no longer guarantee and a fraud it can no longer detect by software, the sector is falling back on the only defences that have ever really worked: putting a human in the room and taking the machine out of it. The most symbolically loaded of these retreats happened at Princeton.

In May 2026, Princeton's faculty voted, with a single dissenting voice, to require that all in-person examinations be supervised by instructional staff. That may sound like housekeeping. It was not. Princeton had operated since 1893 on an honour code under which students sat their examinations unproctored, pledging in writing not to cheat and, crucially, accepting a collective responsibility to report classmates who did. For a hundred and thirty-three years, no invigilator stood at the front of a Princeton exam hall. That tradition of unsupervised examination is now over. The new policy, which took effect on 1 July 2026, places a proctor in every room, present as what the faculty legislation described as a witness rather than an enforcer, but a witness nonetheless — an institutional admission that the honour system, as a mechanism for guaranteeing integrity, has failed.

The reasons the faculty gave are as revealing as the vote itself. Reporting by The Daily Princetonian, which broke the story, drew on a 2025 survey of some five hundred graduating seniors, thirty per cent of whom admitted to having cheated at least once. But the more telling finding concerned the enforcement mechanism rather than the offence. Students said they found it increasingly difficult to identify cheating in a modern exam hall — a phone under the desk is far harder to spot than a crib sheet up a sleeve — and, more damningly, that they were unwilling to inform on their peers for fear of social retaliation, of being teased, doxxed or ostracised. The survey put a figure on that reluctance which is difficult to argue away. Nearly forty-five per cent of the seniors who responded knew of honour code violations that they had chosen not to report; just four in every thousand had ever reported a peer. An honour code depends on students being both able and willing to police one another. Princeton's students, by their own testimony, had become neither. The code had become a ritual with no engine behind it, a signature on a form that certified nothing, and the faculty finally voted to stop pretending otherwise.

Beneath the symbolic drama of Princeton, a quieter and more practical counter-revolution has been under way across the sector, and it has a distinctly analogue flavour. The blue book — the flimsy stapled booklet of lined paper in which generations of students once scrawled their handwritten answers under the eye of an invigilator — is enjoying an improbable renaissance. Sales at the University of Florida reportedly rose by half over two academic years; at the University of California, Berkeley, they were said to be up by four-fifths across the same period, and at Texas A&M by about thirty per cent. Handwriting, it turns out, is one of the few technologies that reliably locks the machine out of the room. You cannot prompt a chatbot with a pen. Alongside the blue books, institutions are experimenting with the oral examination, the ancient viva voce in which a student must explain and defend their reasoning aloud, in real time, to an examiner who can ask follow-up questions the student had no way to prepare with software. The economics department at Brown, in the immediate aftermath of Serrano's discovery, publicly concluded that the in-person examination was the way forward.

None of this is confined to the American Ivy League, and in Britain the retreat has begun to acquire a regulatory edge that the American version still lacks. On 17 August 2026, the think tank Policy Exchange published a report by Philip M. Newton, a professor at Swansea University Medical School, under a title that reads almost as a summary of everything argued here: Evidence of Learning Through Assessment: Protecting the Value of a Degree in the Age of AI. Its central contention is regulatory rather than pedagogical. Summative examinations sat remotely and without supervision, Newton argues, “completely, and obviously, fail” the Office for Students' condition B4, the requirement that assessment be a valid and reliable measure of what a student actually knows. His recommendation is not that universities reform such examinations but that they stop setting them at once, and that the Office for Students and the Quality Assurance Agency compel them to do so if they will not.

The evidence he assembled through freedom of information requests to British universities is the more startling half of the document, because it establishes that the practice is not marginal. In 2023-24, seventy-eight per cent of UK universities used remote online examinations for summative assessment. Only around one in ten invigilated all of them. Two-thirds of the institutional policies governing those examinations made no mention of generative AI whatsoever, well over a year after ChatGPT's release. And seventy per cent of the universities surveyed intended to carry on with unsupervised online examinations regardless. That last figure is worth sitting with. It is not a portrait of a sector caught unawares by a fast-moving technology. It is a portrait of a sector that has been told what is happening, has measured it, and has decided to continue.

Newton's proposed remedies rhyme with what Princeton and Brown have arrived at by harder experience — most obviously a far greater use of the interactive oral viva, on the model much of continental Europe never abandoned — but one of them addresses the signalling problem head-on rather than obliquely. If the certificate can no longer be trusted on its own, he suggests, then the transcript should carry more information about how the student was assessed: not merely the marks earned, but the conditions under which they were earned. It is a modest proposal with immodest implications, because it concedes that the single undifferentiated signal is broken and sets out to replace it with a granular one, in which an invigilated first is legible as something different from an unsupervised one. Employers would learn to read the distinction quickly enough. So, rather less comfortably for the institutions, would applicants choosing between them.

Who actually owns the mess

There is a natural instinct, watching all this, to reach for the language of individual morality — to say that the students cheated, that cheating is wrong, and that the responsibility therefore rests with them. That is true as far as it goes, and it does not go very far, because it explains a scandal in one classroom while leaving the systemic collapse entirely unaccounted for. Twenty-two individual moral failures do not produce a class average of ninety-six on a hardened exam. Something larger is malfunctioning, and the responsibility for it is distributed far more widely than the students who happened to get caught.

The students bear the most immediate and personal responsibility, and it would be sentimental to pretend otherwise. They signed an honour code. They understood, because everybody understands, that submitting a machine's work as their own is a form of fraud. Nobody prompted a chatbot into producing a proof by accident. But it is worth being precise about the incentive structure they were operating inside, because it was engineered, over decades, to reward exactly the behaviour it now punishes. These are young people who were selected, drilled and admitted on the basis of their capacity to optimise every measurable metric of achievement — to treat the grade as the goal and the learning as the incidental means. An admissions arms race that rewards the maximisation of scores above all else should not be astonished when the students it selects go on to maximise their scores by the most efficient means available. The machine is simply the most efficient means yet invented.

The institutions bear a heavier and more culpable share, because they have known about this for three years and have, for the most part, temporised. Generative AI capable of answering undergraduate problem sets has been freely available since late 2022. In the time since, universities have issued a great deal of guidance, convened a great many committees, purchased a great deal of unreliable detection software, and changed astonishingly little about the fundamental architecture of how they assess and certify their students. The take-home essay, the unproctored problem set, the online quiz — the assessment formats most trivially defeated by a chatbot — remained in widespread use long after it was obvious to anyone paying attention that they had become meaningless. Princeton's vote and Brown's pivot to in-person finals are welcome, but they arrived years into an emergency that any honest observer could see coming from the moment ChatGPT launched. The institutions were slow because acting quickly was expensive and disruptive and politically awkward, and because the debasement of the credential is a slow-motion catastrophe that never quite forces a reckoning in any single quarter. They protected their short-term convenience at the expense of the long-term value of the very thing they exist to sell.

Brown's own conduct once Serrano brought his evidence forward is a compact illustration of the reflex. He submitted his findings to the university's Standing Committee on the Academic Code on 16 May 2026, and heard nothing back. Six weeks of silence later, at the end of June, he told the story to the Spanish newspaper El País, and it travelled around the world within days. The committee, mute through May and most of June, made contact shortly after publication, asking him to file individual complaints against each suspected student and to supply copies of their scripts. He provided the additional material on 8 July, at which point a formal investigation was opened and the students began to be contacted one by one. Serrano's verdict on that sequence is unsparing. “It's absolutely clear to me that if I hadn't gone public, nothing would have happened,” he said. He had already described the university's response as meek, and reported that it was seen as appalling and insufficient by the hundreds of people who had written to him in support, many of them Brown alumni.

Brown disputes the characterisation, and its account deserves a hearing rather than a dismissal. The university maintains that it treats every allegation of academic dishonesty with the utmost seriousness, that multiple academic leaders were in contact with Serrano during May about how the allegations could be formally adjudicated, and that the standing committee could not proceed until he supplied the particular details its procedures require — which he did on 8 July, whereupon it moved. Serrano himself later said he was appreciative of Brown finally looking at the case. Both accounts can be true simultaneously, and the fact that they can is precisely the point. A process that is procedurally impeccable and glacially slow is not a defence of the credential; it is a description of how a credential is debased, one unhurried committee cycle at a time. The sanctions available under Brown's academic code run from reprimand to expulsion. As of late August 2026, months after forty perfect scores appeared on a deliberately hardened examination, no outcome has been disclosed and the matter remains unresolved.

And the technology companies bear a share too, though they are the least willing to admit it and the least likely to be held to account. They released, into an education system built entirely around the assumption that a student's submitted work reflects the student's own effort, a tool that shattered that assumption overnight, and they did so with no serious mechanism for allowing institutions to distinguish their product's output from human work. The watermarking that might have made co-existence possible was deprioritised or abandoned because reliable watermarking is commercially inconvenient — it makes your product easier to police and therefore less attractive to precisely the users who most want to hide their tracks. The externality was dumped, as externalities usually are, on someone else's balance sheet. In this case the balance sheet belongs to every institution whose credential now certifies less than it did, and to every honest student whose genuine degree is now shadowed by the suspicion that it might have been faked.

What restoring trust would actually cost

If the value of an élite degree rests, as the economists insist, on its being a costly and hard-to-forge signal, then restoring that value requires making the signal costly and hard to forge again. There is no clever software that does this. There is no policy memorandum that does this. There is only the unglamorous, expensive work of rebuilding assessment around conditions a machine cannot infiltrate, and accepting the price that comes with it.

That price is real, and it is worth naming honestly rather than pretending the return to proctors and blue books is cost-free. Supervised, handwritten, oral and in-person assessment is more labour-intensive, more expensive, less scalable and less accessible than the frictionless online formats it replaces. It disadvantages the student with a disability who needs accommodation, the student whose handwriting cannot keep pace with their thinking, the student for whom a high-pressure oral examination is a crueler test of nerve than of knowledge. The take-home exam that Serrano offered was, remember, an act of compassion — an attempt to accommodate genuine trauma. The move back towards the invigilated hall is a move back towards a harsher, less forgiving, less flexible model of assessment, and the students who will pay the highest price for it are not the confident cheats but the vulnerable and the anxious. That is the bitter irony threaded through the whole affair. The cheating of the many is purchasing a harsher regime for the honest, and the compassion that made the fraud possible will be among its first casualties.

But the alternative to paying that price is worse, because the alternative is a credential that certifies nothing, and a credential that certifies nothing is not merely worthless — it is actively corrosive. It devalues the qualification of every honest graduate retroactively. It corrodes the trust of every employer, every professional body, every graduate admissions committee that has to decide whether the paper in front of them means what it says. It hollows out, from the inside, the single most valuable asset that an institution like Brown or Princeton or Harvard possesses, which is not its endowment or its buildings or its faculty but the simple, centuries-in-the-making public confidence that its name on a certificate is a guarantee of something. That confidence is far easier to destroy than to rebuild. It is the accumulated deposit of generations of credible assessment, and it can be spent down to nothing in the span of a few cohorts who were allowed to fake the signal because catching them was inconvenient.

What the Brown scandal ultimately exposes is that this confidence was resting on far more fragile foundations than anyone cared to admit. The whole edifice depended on a detection capability that, it turns out, amounts to little more than a numerate professor with a suspicious mind and a spare afternoon to interrogate his own exam with a chatbot. That is not a system. It is a happy accident that will not recur reliably, and it caught only the careless. The genuinely able cheat — the one who paraphrases, who varies the prompt, who introduces a few deliberate imperfections — was never in any danger, and remains in no danger now. The uncomfortable implication is that the degrees being conferred this summer, at Brown and everywhere like it, carry a guarantee that the institutions issuing them can no longer actually make good on. They are certifying, in many cases, they know not what.

Serrano's question, stripped of its anger, is exactly the right one, and it deserves to be asked not rhetorically but institutionally, by every provost and dean and examinations board in the sector. If the work can be done by pressing a button, what is the degree for? The answer cannot be nothing, because a great many people — employers, governments, students who genuinely learned something, societies that need their doctors and engineers and economists to actually know things — depend on the answer being something. But arriving at a defensible answer will require the institutions to do what they have spent three years avoiding: to accept that the frictionless, scalable, trusting model of assessment they had grown comfortable with is finished, and to pay, in money and labour and lost convenience, for the harder and more human forms of examination that can still tell the wheat from the chaff. The bill for restoring the signal has come due. The only remaining question is whether the universities will settle it now, while there is still a signal left to save, or keep temporising until the currency they print is worth no more than the paper it is printed on.

References

  1. Preston Fore, “'Humanity has chosen to become idiots': This Brown professor switched to take-home exams after a mass shooting and discovered mass cheating,” Fortune, 29 June 2026. https://fortune.com/2026/06/29/roberto-serrano-brown-university-massacre-ai-cheating/
  2. “AI-Driven Cheating Scandal Uncovered at Brown University,” OECD.AI Incidents, 28 June 2026. https://oecd.ai/en/incidents/2026-06-28-3184
  3. “Brown Professor Suspects Most of His Class Used AI to Cheat,” Inside Higher Ed, 8 July 2026. https://www.insidehighered.com/news/faculty/learning-assessment/2026/07/08/brown-professor-suspects-most-his-class-used-ai-cheat
  4. “Brown University professor raises AI cheating concerns,” The Boston Globe, 15 July 2026. https://www.bostonglobe.com/2026/07/15/metro/brown-university-ai-suspected-cheating/
  5. “After AI cheating concerns, economics professors see in-person exams as a path forward,” The Brown Daily Herald, April 2026. https://www.browndailyherald.com/article/2026/04/after-ai-cheating-concerns-economics-professors-see-in-person-exams-as-a-path-forward
  6. “Princeton Introduces Proctoring, Changing Honor Code,” Inside Higher Ed, 15 May 2026. https://www.insidehighered.com/news/faculty/learning-assessment/2026/05/15/princeton-introduces-proctoring-changing-honor-code
  7. “Princeton faculty mandate proctoring for in-person exams, upending 133 years of precedent,” The Daily Princetonian, May 2026. https://www.dailyprincetonian.com/article/2026/05/princeton-news-adpol-proctoring-in-person-examinations-passed-faculty-133-years-precedent
  8. “The Graduating Class of 2024 By the Numbers — Academics,” The Harvard Crimson, 2024. https://features.thecrimson.com/2024/senior-survey/academics/
  9. “The Graduating Class of 2025 By the Numbers — Academics,” The Harvard Crimson, 2025. https://features.thecrimson.com/2025/senior-survey/academics/
  10. “The Graduating Class of 2026 By the Numbers — Academics,” The Harvard Crimson, 2026. https://features.thecrimson.com/2026/senior-survey/academics/
  11. “47% of Harvard seniors admit to cheating — and the problem existed long before ChatGPT,” Fortune, 23 June 2026. https://fortune.com/2026/06/23/harvard-cheating-academic-integrity-ai-detection/
  12. Zach Hrynowski and Stephanie Marken, “AI Is Routine for College Students, Despite Campus Limits,” Gallup (Lumina Foundation-Gallup 2026 State of Higher Education study), 2026. https://news.gallup.com/poll/704090/routine-college-students-despite-campus-limits.aspx
  13. “Majority of college students use AI for their coursework, poll finds,” UPI, 2 April 2026. https://www.upi.com/Top_News/US/2026/04/02/survey-college-students-artificial-intelligence-coursework/5341775162201/
  14. “The Truth About Turnitin's AI Detection Accuracy in 2025,” Turnitin, 2025. https://turnitin.app/blog/The-Truth-About-Turnitins-AI-Detection-Accuracy-in-2025.html
  15. “Limitations of AI Detection Tools,” Artificial Intelligence Steering Council, Brandeis University, 2025. https://www.brandeis.edu/ai-steering-council/ai-literacy/ai-teaching-learning/detection-tools.html
  16. “Generative AI Detection Tools — The Problems with AI Detectors: False Positives and False Negatives,” Legal Research Center, University of San Diego, 2025. https://lawlibguides.sandiego.edu/c.php?g=1443311&p=10721367
  17. “Survey on Plagiarism Detection in Large Language Models: The Impact of ChatGPT and Gemini on Academic Integrity,” arXiv, 2024. https://arxiv.org/pdf/2407.13105
  18. Michael Spence, “Job Market Signaling,” The Quarterly Journal of Economics, Vol. 87, No. 3, 1973. https://doi.org/10.2307/1882010
  19. Bryan Caplan, The Case Against Education: Why the Education System Is a Waste of Time and Money, Princeton University Press, 2018. https://press.princeton.edu/books/hardcover/9780691174655/the-case-against-education
  20. “Schools fight AI cheating with return to pen and paper blue books,” Fox News, 2026. https://www.foxnews.com/tech/schools-turn-handwritten-exams-ai-cheating-surges
  21. “Blue books are back: The revival of pen and paper exams,” The Daily Cardinal, 6 November 2025. https://www.dailycardinal.com/article/2025/11/blue-books-are-back-the-revival-of-pen-and-paper-exams
  22. “Colleges Turn to Oral and Handwritten Exams as AI Disrupts Assessments,” eWEEK, 2026. https://www.eweek.com/news/colleges-turn-to-oral-exams-ai-disruption/
  23. “Are universities returning to in-person exams to combat AI cheating?,” Times Higher Education, 2025. https://www.timeshighereducation.com/depth/are-universities-returning-person-exams-combat-ai-cheating
  24. “Ban all remote unsupervised tests 'immediately', urges report,” Times Higher Education, 18 August 2026. https://www.timeshighereducation.com/news/ban-all-remote-unsupervised-tests-immediately-urges-report
  25. Philip M. Newton, Evidence of Learning Through Assessment: Protecting the Value of a Degree in the Age of AI, Policy Exchange, 17 August 2026. https://policyexchange.org.uk/publication/evidence-of-learning-through-assessment/

Tim Green

Tim Green UK-based Systems Theorist & Independent Technology Writer

Tim explores the intersections of artificial intelligence, decentralised cognition, and posthuman ethics. His work, published at smarterarticles.co.uk, challenges dominant narratives of technological progress while proposing interdisciplinary frameworks for collective intelligence and digital stewardship.

His writing has been featured on Ground News and shared by independent researchers across both academic and technological communities.

ORCID: 0009-0002-0156-9795 Email: tim@smarterarticles.co.uk

Listen to the free weekly SmarterArticles Podcast

Discuss...

Six thousand office workers in the United States, the United Kingdom and Australia were asked, across December 2025 and January 2026, how much time artificial intelligence had given back to them. They said eleven hours a week.

Then they were asked a second question, the one almost nobody asks. How much time do you spend feeding these systems context, supervising what comes out, debugging the mistakes and cleaning up afterwards? The answer was 6.4 hours a week.

That is the whole argument in two numbers. Roughly six of every ten hours artificial intelligence appears to save are consumed by the labour of making artificial intelligence work. The survey, published in June 2026 as the Work AI Index by Glean's Work AI Institute alongside researchers from Stanford, Notre Dame, Emory, UC Berkeley, UC Santa Barbara, UNC Charlotte and University College London, gave the residual a name that has since escaped into general use: botsitting. Rebecca Hinds, who heads the institute and co-authored the report, and her colleagues defined it coldly, as “the largely unrecognized, unbudgeted, and untracked labor of making AI usable”.

Unrecognised, unbudgeted, untracked. Three adjectives doing an enormous amount of work.

On 18 August 2026, the Indian human resources publication HRKatha ran an analysis of the same phenomenon under the more homely label of bot sitting, defining it as the work employees do prompting systems, checking answers, correcting errors, supplying missing context, trying again when outputs go wrong, and deciding whether the eventual result can be trusted at all. Its sharpest line concerns measurement. The real test of efficiency, the piece argued, is not how quickly a machine produces an answer but how much human work remains before anyone is willing to trust it.

Nobody is measuring that. And what nobody measures, nobody pays for.

The Arithmetic That Stops Working at Five Minutes

Take the scenario in its simplest form. A task used to take twenty minutes of a person's own work. Now it takes five minutes of generation followed by fifteen minutes of checking, correcting and verifying. On the stopwatch, nothing has changed. On the productivity dashboard, everything has. The dashboard records five minutes of task completion, because that is the interval in which the tool was invoked and the output produced. The fifteen minutes afterwards are logged as the worker doing their job, which is what they were doing before the tool arrived.

The tool has not saved fifteen minutes. It has reclassified them.

This is a straightforward consequence of how enterprise software reports on itself. Adoption metrics count seats, prompts and sessions. They do not count the second and third attempts, the cross-check against a source document, the quiet decision not to send the thing at all. More prompts is not a measure of value. It may be a measure of the opposite: a worker who prompts a system eleven times before getting a usable answer generates eleven data points that a usage dashboard reads as enthusiasm.

The Work AI Index found a second figure that should worry anyone relying on those dashboards. Only thirteen per cent of organisations surveyed said artificial intelligence had significantly improved their performance, against eighty-seven per cent of workers who said they were using it. That gap between individual time saved and organisational performance gained is where the botsitting hours have gone. They have not disappeared. They have been absorbed into a category of labour the accounting system cannot see.

The survey also documented what happens when supervision becomes unaffordable. Sixty-nine per cent admitted to what the report calls botshitting: shipping output they had not verified, did not fully understand, or could not confidently stand behind. Forty-one per cent had sent work they would be unable to explain if questioned. Twenty-eight per cent admitted blaming the machine for their own mistakes. Workers who reported botshitting were 3.8 times more likely to be looking for another job.

That last figure is the tell. This is not laziness. It is triage under an unfunded mandate.

Lisanne Bainbridge Wrote the Manual for This in 1983

None of this is new. The person who explained it most economically did so forty-three years ago, in a five-page paper about process control in factories and power stations.

Lisanne Bainbridge published “Ironies of Automation” in the journal Automatica in 1983. Her central observation was that the designer of an automated system regards the human operator as unreliable and inefficient, and so tries to design them out, but cannot automate everything. What remains for the human is precisely the residue that could not be specified: the awkward, ambiguous, judgement-heavy fragments that defeated the engineering. The operator is left with the hardest parts of the job, stripped of the easier parts that used to keep their skills sharp, and asked to intervene in exactly the situations for which they are now least prepared.

Bainbridge was blunt about monitoring in particular. “The human monitor has been given an impossible task,” she wrote, noting that where a computer is making decisions faster and on more dimensions than a person can follow, “there is therefore no way in which the human operator can check in real-time that the computer is following its rules correctly.” She caught the training paradox too: it is ironic, she observed, to train operators in following instructions and then place them in the system to provide intelligence.

Substitute a large language model for a distributed control system and the paper reads like a memo from last week. The knowledge worker of 2026 has been handed the residue. Drafting a first version of a summary, a function, a customer reply: that was the tractable part, and it has been automated. What is left is knowing whether the draft is right, which requires knowing the domain, which was previously maintained by doing the tractable part.

The literature that followed quantified the failure modes she predicted. Raja Parasuraman and Dietrich Manzey's 2010 review in Human Factors, synthesising decades of empirical work on automation complacency and automation bias, reached conclusions that should be printed on the login screen of every enterprise assistant. Complacency emerges specifically under multiple-task load, when manual tasks compete with the automated task for attention. It appears in expert users as readily as in novices, and cannot be trained away with simple practice. Automation bias produces both errors of omission, where the person fails to act because the system did not flag a problem, and errors of commission, where the person acts wrongly because the system told them to. Neither is reliably prevented by instructions telling people to be careful.

Meanwhile the monitoring itself is not free. Joel Warm, Parasuraman and Gerald Matthews titled their 2008 Human Factors paper on the subject with unusual directness: vigilance requires hard mental work and is stressful. Sustained attention to a mostly reliable process is not a restful state between bouts of real work. It is a demanding task with measurable workload and stress costs, and performance on it degrades over time.

So the corporate framing, in which the human is elevated from doing to overseeing, describes a promotion. The human factors literature describes a transfer into a job that is cognitively expensive, psychologically taxing, unavoidably error-prone and, in most workplaces, entirely uncompensated.

Sixteen Developers Who Were Certain They Had Gone Faster

The most instructive evidence in this debate is a small randomised controlled trial with an awkward result and an unusually honest set of authors.

In 2025, the research organisation METR recruited sixteen experienced open-source developers working on repositories they personally maintained, projects averaging over 22,000 GitHub stars. It randomised 246 real issues from those repositories into two conditions: artificial intelligence tools permitted, or not permitted. Tasks averaged around two hours. Beforehand, the developers forecast that the tools would speed them up by twenty-four per cent.

They were nineteen per cent slower with the tools.

What happened next matters more than the slowdown. After completing the work, having lived through the actual elapsed time, the same developers estimated that artificial intelligence had made them twenty per cent faster. They were wrong by roughly forty percentage points about their own labour, in the direction of the tool, on tasks they had personally performed within the previous few hours.

That is the epistemological problem at the heart of every self-reported productivity statistic in this field, including the eleven hours in the Work AI Index. Generation feels fast because it is fast and visible. Verification feels like ordinary work because it is ordinary work, diffuse, arriving in fragments scattered through the day. People are demonstrably bad at summing the second category and comparing it to the first.

METR deserves credit for what it did afterwards. In February 2026 it announced it was redesigning the experiment. Its follow-up cohorts produced point estimates of a negative eighteen per cent speedup for the original developers, with a confidence interval from negative thirty-eight to positive nine per cent, and negative four per cent for newly recruited developers. The slowdown persisted in the point estimates but the intervals now straddled zero. More importantly, METR reported that developers were increasingly refusing to take part in conditions barring them from using the tools, that the pay rate had fallen from 150 dollars an hour to fifty, and that measuring elapsed time had become genuinely hard when participants ran multiple agents concurrently. “Due to the severity of these selection effects, we are working on changes to the design of our study,” the organisation wrote, adding that the true speedup among excluded developers could be considerably higher.

That is what intellectual honesty looks like, and it cuts both ways. Anyone citing the nineteen per cent slowdown as a settled fact about artificial intelligence in 2026 is overreaching. But the perception gap is the more durable finding, and nothing in the update disturbs it. The difficulty in the follow-up was not that developers had become good at estimating their own throughput. It was that the experiment could no longer isolate the variable, because people who use these tools will no longer agree to stop.

Why Checking Is Harder Than Doing

There is a structural reason verification consumes more time than intuition suggests, and it concerns the shape of machine error.

A junior colleague who does not know something produces work that signals its own uncertainty. The prose is hedged, the gaps are obvious, the citations are missing. A language model that does not know something produces work that is fluent, confident, correctly formatted and internally consistent. The error sits in a sentence that looks exactly like every true sentence around it. The cost of finding it is therefore not the cost of scanning for anomalies. It is the cost of independently establishing the truth of each load-bearing claim, which in the limit is the cost of having done the work yourself.

This is why the five-minutes-plus-fifteen arithmetic is not a transitional inconvenience that better models will erase. As accuracy rises, the frequency of error falls but the difficulty of detection rises, because a rarer error in more plausible packaging demands more sustained vigilance to catch. Parasuraman and Manzey's finding is precisely that reliable automation breeds the attentional withdrawal making occasional failure catastrophic. Better models make botsitting less frequent and more consequential at once.

The downstream costs have now been measured in money. In September 2025, BetterUp Labs and the Stanford Social Media Lab surveyed 1,004 full-time American desk workers about what they termed workslop: output with the appearance of good work but lacking the substance to advance the task. Forty per cent had received it in the preceding month, and respondents estimated that 15.4 per cent of the work reaching them fell into the category. Each incident took an average of one hour and fifty-one minutes to sort out, around twenty minutes longer than doing the work properly in the first place would have taken. Converted using respondents' own salaries, that came to roughly 186 dollars per employee per month, or more than nine million dollars a year for an organisation of ten thousand people. Managers were markedly more exposed than individual contributors, at fifty-four per cent against 38.5 per cent.

Note what workslop is in labour terms. It is botsitting skipped upstream, landing unpriced on somebody downstream. The person who declined to verify saved fifteen minutes. The person who received the output spent one hour and fifty-one. That is not productivity. It is a transfer of unpaid supervisory work between colleagues, with interest.

What Gets Measured Was Never Designed to See This

Organisations do have instruments for measuring cognitive workload. They simply do not use them for this.

The NASA Task Load Index, developed by Sandra Hart and Lowell Staveland in 1988 and still the most widely used subjective workload instrument in ergonomics, decomposes workload into six components: mental demand, physical demand, temporal demand, performance, effort and frustration. In its original form it asks respondents to weight those components against one another across all fifteen possible pairs. Hart's twenty-year retrospective in 2006 catalogued its spread across aviation, medicine and interface design.

Look at that list and notice how badly conventional workload accounting maps onto botsitting. Corporate measurement, where it exists at all, tracks temporal demand: hours worked, tickets closed, tasks completed. Mental demand, effort and frustration are exactly the dimensions supervision loads most heavily and timesheets record least. A worker who closes the same number of tickets while spending two extra hours a day in a state of alert scepticism about machine output registers, on every dashboard the employer owns, as having had an identical day.

This is where the sociology of work becomes more useful than the economics. Susan Leigh Star and Anselm Strauss published “Layers of Silence, Arenas of Voice” in the journal Computer Supported Cooperative Work in 1999, a foundational paper on invisible work in computer-supported systems. Their opening move was to insist that what counts as work is a matter of definition, not a matter of fact, and that designers routinely automate the visible portion of a process while remaining oblivious to the articulation work holding the arrangement together: the coordinating, contextualising, patching and repairing that never appears in a process diagram and therefore never appears in a budget.

Botsitting is articulation work. Supplying missing context to a model, deciding which of three plausible outputs matches what the client meant, knowing the tool has silently used a superseded policy document: this is the connective labour Star and Strauss described, now performed on behalf of a machine rather than a colleague. Their warning was that automating around invisible work does not eliminate it. It removes the vocabulary for discussing it.

The lifecycle synthesis of human-AI collaboration risks published on arXiv in August 2026 by Md Foysal Ahmed, Isaac Kobby Anni and Md Main Uddin Rony makes a compatible point in the language of risk taxonomy. Surveying evidence across healthcare, journalism, education, research, organisational decision-making and defence, the authors identify six recurring risk clusters that cut across domains: trust miscalibration, cognitive burden, accountability gap, capability erosion, goal misalignment, and AI anxiety and technostress. These are not separable engineering defects to be fixed one at a time, they argue, but interlocking sociotechnical dynamics cascading through four lifecycle stages, which is why piecemeal interventions frequently create unintended consequences. Cognitive burden and accountability gap sitting adjacent in the same taxonomy is no coincidence. They are one problem seen from the worker's side and the organisation's side.

The History of Labour-Saving Devices Is a History of Redistribution

There is a precedent for all of this, and it is not from computing.

Ruth Schwartz Cowan's “More Work for Mother”, published in 1983 and awarded the Society for the History of Technology's Dexter Prize the following year, asked why American housewives were working longer hours in 1970 than in 1870 despite a century of mechanisation. Washing machines, vacuum cleaners, gas ovens, commercial flour and refrigeration had all arrived. The hours had not fallen.

Cowan's answer was that the appliances did not eliminate labour. They redistributed it and then raised the standard it was held to. Work previously done by servants, husbands, children and commercial services was pulled back onto one person, while expectations for cleanliness, nutritional variety and laundry frequency rose to consume whatever slack the machines created. The technology was genuinely labour-saving per unit of output. Total labour rose anyway, because output expectations rose faster and the residual tasks landed on a single unpaid worker.

Read the Work AI Index against that and the shape is familiar. Eleven hours of unit-level saving, 6.4 hours of new residual labour, thirteen per cent of organisations reporting real improvement, and a rising expectation of volume absorbing the difference.

The digital version of this dynamic already has a well-documented lower tier. Mary L. Gray and Siddharth Suri's “Ghost Work”, published in 2019, described the invisible global workforce that fills the gaps automated systems cannot close: content flagging, transcription checking, data labelling, the human intervention that makes a service look seamless. Their formulation of the paradox is that the drive to eliminate human labour reliably generates new human tasks, and that those tasks are systematically hidden, poorly paid and structurally insecure. The most cited illustration remains Billy Perrigo's January 2023 investigation for TIME, which documented that OpenAI had used workers in Kenya, employed through the outsourcing firm Sama, to label descriptions of child sexual abuse, torture, self-harm and bestiality so that ChatGPT could learn to filter such material. They were paid between 1.32 and two dollars an hour.

The point for the office worker in 2026 is not that their situation is equivalent. It plainly is not. The point is that the industry has an established pattern of relying on human labour it declines to name, and that the pattern has migrated from the outsourced periphery to the salaried core. The mechanism is identical: apparent autonomy produced by human effort the accounting treats as external to the system.

Article 14 Turns You Into the Accountability Sink

Regulation has made the informal expectation of supervision into a legal duty, and in doing so has clarified who carries the risk.

Article 14 of the European Union's Artificial Intelligence Act requires that high-risk systems be designed so they “can be effectively overseen by natural persons during the period in which they are in use”. The overseeing person must be enabled to “properly understand the relevant capacities and limitations of the high-risk AI system and be able to duly monitor its operation”; to “remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias)”; to decide “not to use the high-risk AI system or to otherwise disregard, override or reverse” its output; and to interrupt the system through a stop button.

Read those clauses as a job description rather than a compliance obligation. They specify a role requiring domain expertise sufficient to override a machine, metacognitive awareness of one's own susceptibility to automation bias, and sustained vigilance across the operational life of the system. There is no corresponding requirement anywhere in the Act that this role be staffed, budgeted, timetabled or paid.

The timing is instructive. Those Article 14 obligations for standalone high-risk systems were originally due on 2 August 2026. The Digital Omnibus deferring them, agreed on 6 May 2026 and confirmed by member state representatives a week later, was published in the Official Journal on 24 July 2026 and entered into force on 27 July, six days before the deadline it displaced. The revised dates are settled law: 2 December 2027 for standalone high-risk systems, and 2 August 2028 for those embedded in products already covered by European Union product safety law. The Article 50 transparency duties stayed on the original schedule and took effect earlier this month.

So the expectation that a human will supervise the machine is already operating inside every workplace that has deployed one, and the legal duty to build machines that can actually be supervised has been postponed by sixteen months. That ordering places the burden on the human before it places the corresponding design obligation on the system.

Ben Green, in a 2022 paper in Computer Law and Security Review, surveyed forty-one policies mandating human oversight of government algorithms and found two connected flaws. The evidence suggests people cannot reliably perform the oversight functions the policies assume, and as a result the policies legitimise deployment of faulty systems without addressing what is wrong with them. His remedy was to shift accountability from individual oversight to institutional oversight.

Madeleine Clare Elish gave the failure mode its name. In “Moral Crumple Zones”, published in Engaging Science, Technology, and Society in 2019, she analysed accidents involving complex automated systems and observed that responsibility is routinely misattributed to the human closest to the failure, even where that human had minimal control over the system's behaviour. The crumple zone in a car absorbs impact to protect the occupant. The moral crumple zone absorbs blame to protect the integrity of the technological system, at the expense of the nearest operator.

Twenty-eight per cent of Work AI Index respondents admitted blaming artificial intelligence for their own mistakes. Elish's argument is that the institutional traffic runs overwhelmingly the other way.

The Skills You Stop Using Are the Skills You Need to Check With

Bainbridge's most uncomfortable prediction was that operators would lose the competence they needed for the interventions automation reserved for them. Medicine has now produced the cleanest demonstration.

In October 2025, The Lancet Gastroenterology and Hepatology published a multicentre observational study of endoscopist deskilling drawn from four Polish centres in the ACCEPT trial, which introduced computer-aided polyp detection at the end of 2021 and then randomised subsequent colonoscopies to proceed with or without artificial intelligence assistance. Nineteen experienced endoscopists, each with more than two thousand colonoscopies behind them, were studied. Their adenoma detection rate in unassisted colonoscopies fell from 28.4 per cent before exposure to the tool to 22.4 per cent afterwards, an absolute decline of six percentage points.

These were not trainees. They were highly experienced clinicians whose unaided performance degraded measurably after routine assistance, on a metric directly linked to cancer prevention.

The implication for botsitting is recursive and unpleasant. The supervisory role exists because the human is supposed to catch what the machine gets wrong. That requires the domain judgement previously maintained by performing the task unaided, which is exactly what the tool has removed. The capability erosion cluster in the arXiv risk taxonomy captures the structure: the intervention creating the need for oversight simultaneously degrades the capacity to provide it. Any organisation counting on human verification as its safety net is depending on a resource its own deployment strategy is quietly consuming.

Where the Productivity Actually Landed

The macro evidence is not that artificial intelligence does nothing. It is that the gains are real, narrow, and much smaller in aggregate than the discourse implies.

The strongest firm-level result remains the study by Erik Brynjolfsson, Danielle Li and Lindsey Raymond in the Quarterly Journal of Economics in 2025, examining the staggered rollout of a generative conversational assistant across 5,172 customer support agents. Access to the assistant raised issues resolved per hour by fifteen per cent on average. Crucially, the effect was concentrated among novice and lower-skilled workers, with minimal or slightly negative effects for the most experienced. The tool distributed the accumulated tacit knowledge of the best performers to everyone else.

That result is genuine and it is also specific. Customer support has high task volume, structured interactions, immediate feedback and low per-instance verification cost. It is the best case, and not obviously the case for legal drafting, engineering, medical documentation or financial analysis, where verifying an output can cost more than producing it and the consequences of an unverified error arrive months later.

The national statistics tell a thinner story. The US Bureau of Labor Statistics reported on 6 August 2026 that nonfarm business labour productivity rose at an annualised rate of 1.4 per cent in the second quarter of 2026, and 2.2 per cent measured against the second quarter of 2025, with output up 2.5 per cent and hours worked up 0.2 per cent over the year. Unit labour costs rose at an annualised 1.3 per cent in the quarter. These are respectable figures. They are not the signature of a technology that has removed eleven hours a week from the working lives of eighty-seven per cent of office workers.

The Upwork Research Institute, surveying around 2,500 people across the United States, United Kingdom, Australia and Canada in the spring of 2024, found the gap in its rawest form. Ninety-six per cent of C-suite leaders expected artificial intelligence to raise productivity. Seventy-seven per cent of employees using it said it had increased their workload. Forty-seven per cent did not know how to achieve the productivity gains their employers expected. One in three said they were likely to quit within six months because of burnout. Kelly Monahan, managing director of the institute, framed the conclusion carefully: it is possible for the technology to raise productivity and improve well-being simultaneously, but that outcome requires a fundamental change in how work and talent are organised, not merely a change in tooling.

Two years on, the Work AI Index suggests the reorganisation has not happened. What has happened is that the residual labour acquired a name.

Counting the Hour That Was Never Saved

Naming a thing is a precondition for measuring it, and measuring it is a precondition for paying for it. Here is what measurement would actually involve.

The first step is to stop treating verification as a discretionary activity performed by conscientious individuals and start treating it as a defined task with an estimated duration. That means workload models built on generation time plus verification time, with the second term populated from observation rather than optimism. The concept already exists. Work presented at the 2026 CHI conference on human factors in computing systems by Guangrui Fan, Dandan Liu, Lihu Pan and Rui Zhang constructed a behavioural verification-load index for programmers from observable signals including compile and test failures, code churn, pauses and context switches, and showed across sixty participants that it tracked both subjective burden and correctness. Notably, in that study the tools reduced measured workload and time on task while the verification-load metric still predicted accumulating stress and fatigue over repeated use. Both things can be true. The gains are real and the residual burden is real, and only one of them currently appears in any management report.

The second step is to use an instrument designed for the purpose. The NASA Task Load Index has been in continuous use for nearly four decades, is free, takes minutes to administer, and measures precisely the dimensions supervision loads and timesheets miss. There is no methodological obstacle to running it before and after deployment of a workplace assistant, only an incentive obstacle: the results might contradict the business case.

The third step is architectural. The paper on safe and responsible artificial intelligence agents submitted to arXiv in January 2026 by Edward Cheng, Jeshua Cheng and Alice Siu argues for a three-pillar model grounded in transparency, accountability and trustworthiness, and for staged autonomy achieved through progressive validation rather than immediate full automation, by explicit analogy with the incremental rollout of autonomous driving. Reviewing the state of the art, the authors point to Magentic-UI, an open-source interface platform for human-in-the-loop agentic systems, as an example of oversight embedded through structured, repeatable mechanisms: co-planning, co-tasking, action approval and answer verification. The significance for botsitting is that these are discrete, nameable, loggable events. An oversight step that exists as a defined interaction in software can be counted, timed, staffed and, if anyone chooses, paid for. Oversight existing only as a cultural expectation that someone will check cannot.

The fourth step is to treat oversight capacity as something an organisation builds rather than something it assumes. Yao Xie and Walter Cullen argued in a December 2025 arXiv paper that major ethics guidelines and laws, the EU AI Act explicitly included, call for effective human oversight without defining it as a distinct and developable capacity. Their proposal situates it within a well-being efficacy framework integrating artificial intelligence literacy, ethical discernment and awareness of human needs, on the grounds that people inevitably project desires, fears and interests onto these systems and that oversight therefore requires the competence to examine and, where necessary, restrain problematic demands.

That identifies the right object. Human oversight is not a checkbox and it is not a personality trait. It is a skill that decays without practice, costs energy to exercise, and is currently being demanded at scale from people who have received no training in it and no allowance for it.

The fifth step determines whether any of this reaches a pay packet, and it is contractual rather than technical. The mechanism by which invisible labour has historically become visible is not measurement alone. It is bargaining. Ben Green's argument for shifting from individual to institutional oversight points the same way: the question is not whether a given worker checked carefully enough, but whether the institution deploying the system created conditions under which careful checking was possible.

Which brings the argument back to the twenty minutes. If the task now takes five minutes of generation and fifteen of verification, the honest description is that the technology has changed the composition of the work without changing its duration, and has shifted its character from production, which most people find satisfying, to inspection, which the vigilance literature has shown for decades to be tiring, stressful and prone to exactly the errors it exists to prevent. That might still be worth doing. Verification scales in ways expertise does not, and the customer support evidence shows real gains where the economics line up.

But the hour supposedly saved is not available for redeployment if a substantial fraction of it was never saved. Counting it as though it were is not an optimistic forecast. It is a measurement error, and the people absorbing the difference are the ones holding the mouse at half past six, reading a paragraph for the third time, trying to work out whether the confident sentence in the middle of it is true.


References

  1. Liji Narayan, “Bot sitting: When humans end up babysitting machines,” HRKatha, 18 August 2026. https://www.hrkatha.com/features/hr-pops-features/bot-sitting-when-humans-end-up-babysitting-machines/
  2. Work AI Institute, Glean, “The Work AI Index 2026,” June 2026. https://www.glean.com/work-ai-institute/reports/work-ai-index
  3. METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” 10 July 2025. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
  4. METR, “We are Changing our Developer Productivity Experiment Design,” 24 February 2026. https://metr.org/blog/2026-02-24-uplift-update/
  5. Upwork Inc., “Upwork Study Finds Employee Workloads Rising Despite Increased C-Suite Investment in Artificial Intelligence,” 23 July 2024. https://investors.upwork.com/news-releases/news-release-details/upwork-study-finds-employee-workloads-rising-despite-increased-c
  6. Lisanne Bainbridge, “Ironies of Automation,” Automatica, vol. 19, no. 6, 1983, pp. 775-779. https://www.sciencedirect.com/science/article/abs/pii/0005109883900468
  7. Raja Parasuraman and Dietrich H. Manzey, “Complacency and Bias in Human Use of Automation: An Attentional Integration,” Human Factors, vol. 52, no. 3, June 2010, pp. 381-410. https://journals.sagepub.com/doi/10.1177/0018720810376055
  8. Joel S. Warm, Raja Parasuraman and Gerald Matthews, “Vigilance Requires Hard Mental Work and Is Stressful,” Human Factors, vol. 50, no. 3, June 2008, pp. 433-441. https://journals.sagepub.com/doi/10.1518/001872008X312152
  9. Sandra G. Hart, “NASA-Task Load Index (NASA-TLX); 20 Years Later,” Proceedings of the Human Factors and Ergonomics Society Annual Meeting, vol. 50, no. 9, October 2006, pp. 904-908. https://journals.sagepub.com/doi/10.1177/154193120605000909
  10. Susan Leigh Star and Anselm Strauss, “Layers of Silence, Arenas of Voice: The Ecology of Visible and Invisible Work,” Computer Supported Cooperative Work, vol. 8, nos. 1-2, March 1999, pp. 9-30. https://dl.acm.org/doi/10.1023/A:1008651105359
  11. Ruth Schwartz Cowan, More Work for Mother: The Ironies of Household Technology from the Open Hearth to the Microwave, Basic Books, 1983. https://www.hachettebookgroup.com/titles/ruth-schwartz-cowan/more-work-for-mother/9780465047321/
  12. Mary L. Gray and Siddharth Suri, Ghost Work: How to Stop Silicon Valley from Building a New Global Underclass, Houghton Mifflin Harcourt, 2019. https://ghostwork.info/
  13. Billy Perrigo, “Exclusive: OpenAI Used Kenyan Workers on Less Than $2 Per Hour to Make ChatGPT Less Toxic,” TIME, 18 January 2023. https://time.com/6247678/openai-chatgpt-kenya-workers/
  14. BetterUp Labs and Stanford Social Media Lab, “Workslop: The Hidden Cost of AI-Generated Busywork,” September 2025. https://www.betterup.com/blog/hidden-costs-workslop
  15. Erik Brynjolfsson, Danielle Li and Lindsey R. Raymond, “Generative AI at Work,” The Quarterly Journal of Economics, vol. 140, no. 2, May 2025, pp. 889-942. https://academic.oup.com/qje/article/140/2/889/7990658
  16. “Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study,” The Lancet Gastroenterology and Hepatology, October 2025. https://www.thelancet.com/journals/langas/article/PIIS2468-1253(25)00133-5/abstract
  17. European Union, Regulation (EU) 2024/1689, Article 14: Human Oversight. https://artificialintelligenceact.eu/article/14/
  18. U.S. Bureau of Labor Statistics, “Productivity and Costs, Second Quarter 2026, Preliminary,” 6 August 2026. https://www.bls.gov/news.release/archives/prod2_08062026.htm
  19. Gibson Dunn, “EU AI Act Omnibus Agreement: Postponed High-Risk Deadlines and Other Key Changes,” 2026. https://www.gibsondunn.com/eu-ai-act-omnibus-agreement-postponed-high-risk-deadlines-and-other-key-changes/
  20. Ben Green, “The Flaws of Policies Requiring Human Oversight of Government Algorithms,” Computer Law and Security Review, vol. 45, 2022. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3921216
  21. Madeleine Clare Elish, “Moral Crumple Zones: Cautionary Tales in Human-Robot Interaction,” Engaging Science, Technology, and Society, vol. 5, 2019, pp. 40-60. https://estsjournal.org/index.php/ests/article/view/260
  22. Md Foysal Ahmed, Isaac Kobby Anni and Md Main Uddin Rony, “Toward Resilient Human-AI Collaboration: A Lifecycle Taxonomy of Sociotechnical Risks and Cascading Failures,” arXiv:2608.05614, 6 August 2026. https://arxiv.org/abs/2608.05614
  23. Edward C. Cheng, Jeshua Cheng and Alice Siu, “Toward Safe and Responsible AI Agents: A Three-Pillar Model for Transparency, Accountability, and Trustworthiness,” arXiv:2601.06223, 9 January 2026. https://arxiv.org/abs/2601.06223
  24. Yao Xie and Walter Cullen, “Beyond Procedural Compliance: Human Oversight as a Dimension of Well-being Efficacy in AI Governance,” arXiv:2512.13768, 15 December 2025. https://arxiv.org/abs/2512.13768
  25. Guangrui Fan, Dandan Liu, Lihu Pan and Rui Zhang, “When Help Hurts: Verification Load and Fatigue with AI Coding Assistants,” Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems. https://doi.org/10.1145/3772318.3791176

Tim Green

Tim Green UK-based Systems Theorist & Independent Technology Writer

Tim explores the intersections of artificial intelligence, decentralised cognition, and posthuman ethics. His work, published at smarterarticles.co.uk, challenges dominant narratives of technological progress while proposing interdisciplinary frameworks for collective intelligence and digital stewardship.

His writing has been featured on Ground News and shared by independent researchers across both academic and technological communities.

ORCID: 0009-0002-0156-9795 Email: tim@smarterarticles.co.uk

Listen to the free weekly SmarterArticles Podcast

Discuss...

Picture the first weeks of university, the specific loneliness of them. You have moved into a room in a city you do not know, surrounded by strangers who all seem, on the surface, to have found their people already. The group chats are humming with names you do not recognise. The dining hall is a geometry problem of where to sit. At night, in a narrow bed under an unfamiliar ceiling, the feeling arrives that everyone else received an instruction manual for belonging that you somehow missed. This is not an exotic experience. It is close to universal, and it is precisely the kind of ordinary, corrosive isolation that an entire industry now promises to dissolve with a chat window.

So a team of psychologists decided to test the promise directly. They took nearly three hundred first-semester students, split them into groups, and gave each group a different thing to do for two weeks. One group texted with another human being, a randomly assigned fellow first-year, a total stranger. One group chatted every day with a purpose-built artificial companion engineered to behave like the most supportive friend imaginable, endlessly attentive, warm, and validating. And one group, the control, simply wrote a single sentence about their day in a private journal. At the start and again at the end, everyone completed the same standard questionnaire that clinicians and researchers use to measure how lonely a person feels.

The result is the kind of finding that ought to give a booming industry pause. Only the students who texted a real human being felt measurably less lonely afterwards. The ones who chatted with the AI companion, the product deliberately designed to be the ideal comforting presence, did no better than the ones who wrote a sentence to themselves in a diary. Given every conceivable advantage, engineered specifically to win, the machine drew level with a notebook.

What the experiment actually found

The study was led by Ruo-Ning Li, a doctoral candidate in psychology at the University of British Columbia, working with the UBC happiness researcher Elizabeth Dunn and collaborators at the University of Pennsylvania, among them Dunigan Folk, Abhay Singh and the computer scientist Lyle Ungar. It was published in the Journal of Experimental Social Psychology in the spring of 2026 under a title that states its question plainly: whether a random human peer is better than a highly supportive chatbot at reducing loneliness over time. The design was a randomised controlled trial, the gold standard of experimental evidence, involving 296 students in their first semester at university, the exact population most exposed to the loneliness the study set out to probe.

Participants were assigned to one of three conditions and asked to keep to it for two weeks. In the human condition, each student was paired at random with another first-year and the two exchanged text messages. In the chatbot condition, students conversed daily with a custom generative AI companion the researchers had built and instructed to listen actively, show empathy, and behave as a friendly, positive and supportive presence, checking in and validating feelings. In the control condition, students wrote a single-sentence journal entry each day. Loneliness was measured with the UCLA Loneliness Scale, the most widely used instrument of its kind in psychological research, administered before the intervention began and again on the fifteenth day.

The headline outcome was stark. The students who texted a human peer showed a genuine, statistically meaningful decline in loneliness over the fortnight, on the order of a nine per cent reduction. The students who chatted with the AI companion registered a reduction of roughly two per cent, indistinguishable from the journaling group who received no social contact at all. This was not because the chatbot did nothing. It reliably lifted people's mood in the moment; talking to it made them feel a little better in the short term. But feeling briefly better and becoming less lonely turned out to be two different things, and only human contact delivered the second.

There is a further detail that deserves attention, because it undercuts the most obvious objection. One might assume the machine lost simply because it was cold, robotic, or unconvincing, the familiar clumsiness of a customer-service bot. The opposite was true. The AI companion, by design and in practice, expressed more overt empathy and validation than the human partners did. The humans were ordinary eighteen-year-olds, awkward, distractible, sometimes slow to reply, occasionally unsure what to say. The bot was tireless and effusive. And still the awkward humans won. Whatever the students paired with a stranger were getting, it was not a superior grade of sympathy. It was something else entirely.

Why the chatbot lost when the deck was stacked in its favour

It is worth dwelling on how heavily the experiment was rigged in the machine's favour, because that is what makes the outcome so instructive. This was not a fair fight that the AI narrowly lost. It was a fight the researchers set up specifically to give the AI every advantage, and it still could not win.

Consider the human comparison group. These were not trained counsellors or charismatic extroverts selected for their warmth. They were randomly paired first-year students, people with no particular skill at emotional support, no obligation to be kind, no script. A stranger texting a stranger is a low bar. Anyone who has endured a stilted exchange with someone they have just met knows how thin and effortful such contact can be. The messages were, by any measure, a modest intervention, the sort of thing that costs nothing and requires no infrastructure beyond a phone.

Now consider the machine. It had been purpose-built for this. It never got bored, never had an exam of its own to revise for, never left a message on read for six hours because it was asleep or sulking or busy. It had been explicitly instructed to be maximally supportive, to listen and validate and check in, to perform the role of the perfect friend without any of the friction a real friend brings. If effusive, always-available emotional support were the active ingredient in curing loneliness, the chatbot should have romped home. It had more of that ingredient than the humans could ever muster, delivered on demand at three in the morning.

Instead it tied with a diary. The lead author had, by her own account, expected a different result, anticipating that interacting with the AI might prove roughly as helpful as texting a fellow student. That expectation was reasonable. The industry has spent years and billions of dollars insisting it is true. The experiment simply did not cooperate. And because the trial was deliberately constructed to flatter the machine, its failure cannot be waved away as a matter of the technology not yet being good enough. The chatbot was good. It was warm, responsive and articulate. Being warm, responsive and articulate turned out not to be the thing that mattered.

The strange arithmetic of giving something back

If not empathy, then what? The researchers' explanation is deceptively simple, and it inverts almost everything the marketing of AI companionship assumes. What made the difference, they argue, was not what the students received but what they were able to give.

When a person texts a struggling stranger, the exchange runs in both directions. You are not merely the object of someone else's attention; you are also called upon to attend to them. The other first-year is nervous too, homesick too, unsure too. In responding to that, in offering a word of reassurance or asking how their week has been, the lonely student becomes, briefly, someone who matters to another person, someone whose care has a destination and an effect. The study found that participants were markedly more likely to offer support back when their partner was human than when it was the chatbot. The human condition activated a loop of mutual care. The AI condition did not, because there was no one on the other side who needed anything.

Li put the mechanism directly. When you are talking with a chatbot, she observed, you never have the chance to give something back; human connection has this back and forth of receiving and giving support that makes us feel we matter. That last word is the hinge of the whole finding. Feeling less lonely, it turns out, is not simply a function of being cared for. It depends just as much on having someone to care for, on the sense that your existence registers in another mind, that you could be missed, that your kindness lands somewhere real.

This is not a soft or sentimental claim. It aligns with one of the oldest and most robust ideas in social science. In 1960 the sociologist Alvin Gouldner published a landmark analysis arguing that a norm of reciprocity, the moral expectation that we return the good we receive, is one of the near-universal principal components of human moral codes, present across cultures and foundational to the stability of social life. We are wired, individually and collectively, around exchange. A relationship in which the flow runs only one way, in which one party gives endlessly and the other only takes, is not experienced as a relationship at all. It is experienced as consumption. The AI companion, however solicitous, can only ever be consumed. It has no needs for you to meet, no vulnerability for you to answer, no capacity to actually receive the care you might offer it. You can type “I hope you're okay” to a language model, but there is no one there for whom it could be true.

This reframes the entire proposition of the companion-app industry. These products are sold as devices for receiving support, and at that narrow task they may even succeed; the study found the bot did lift mood in the moment. But loneliness is not a deficit of support received. It is, at least in significant part, a deficit of mattering, of being needed, of having a place in the web of mutual obligation that binds people together. And mattering is precisely the one thing a machine that needs nothing can never provide, no matter how advanced it becomes. This is not a limitation of the current generation of models, to be solved by the next. It is structural. An entity that cannot genuinely be helped cannot let you experience yourself as a helper.

A problem the World Health Organization now calls a threat

The reason any of this reaches beyond a psychology department is the sheer scale of the ailment these products claim to treat. Loneliness has, in the space of a few years, been reclassified from a private melancholy into a public-health emergency, and by some of the most authoritative bodies in global medicine.

In June 2025 the World Health Organization's Commission on Social Connection published its landmark report, bluntly titled as a path from loneliness to social connection, and the numbers it assembled are difficult to absorb. Around one in six people worldwide, the commission estimated, is affected by loneliness. The consequences are not merely emotional. Loneliness and social isolation were linked to an estimated one hundred deaths every hour, more than 871,000 deaths a year across the period studied, through elevated risks of cardiovascular disease, stroke, type 2 diabetes, depression and anxiety. Social isolation, the commission noted, reaches even further into certain groups, affecting up to one in three older adults and as many as one in four adolescents. The report framed social connection not as a lifestyle luxury but as a neglected pillar of health, alongside diet and exercise, and called for a coordinated global response.

This built on an alarm already sounded in the United States. In May 2023 the Surgeon General, Vivek Murthy, issued a formal advisory declaring loneliness and isolation an epidemic and a public-health crisis, reporting that roughly half of American adults had experienced measurable loneliness and that the mortality impact of chronic social disconnection was comparable to smoking as many as fifteen cigarettes a day, and greater than that of obesity or physical inactivity. The advisory laid out, for the first time, a national strategy to rebuild social connection.

Two things follow from this. The first is that the demand the companion-app industry is meeting is genuine and enormous; the loneliness is real, it is widespread, and it is killing people. The second, more uncomfortable point is that a market this large, addressing a need this acute and this painful, is an extraordinary commercial opportunity, and it is being treated as exactly that. When a condition afflicts a sixth of humanity and is officially designated a lethal threat, the incentive to sell a cheap, scalable, always-available remedy is nearly irresistible, whether or not the remedy works. The UBC study is, in effect, a small trial of that remedy against a placebo. The remedy did not beat the placebo.

The hundreds of millions who already talk to machines

It would be one thing if AI companionship were a niche curiosity, a handful of enthusiasts talking to their phones. It is not. It is already one of the most widely adopted categories of consumer software on earth, and its growth is vertical.

The scale is genuinely hard to hold in the mind. China's Xiaoice, one of the earliest and largest social-chatbot platforms, has been reported to serve on the order of 660 million users. Character.AI, the platform most associated with immersive companion conversation, reported in the region of 233 million registered users by early 2026, with its youngest adult cohort, those aged eighteen to twenty-four, spending well over an hour a day on the service across dozens of daily sessions. Replika, one of the first apps to market an AI companion explicitly as a friend or partner, counts users in the tens of millions. Industry trackers estimated that companion applications across the major app stores had surpassed 220 million cumulative downloads by mid-2025, with downloads climbing by close to ninety per cent year on year. One qualifier is worth entering, because registered totals are the industry's most flattering metric rather than its most honest: a sign-up is not a habit. Character.AI's monthly active users are reckoned to have peaked at around 28 million in mid-2024 and drifted back to roughly 20 million by early 2026, so at the level of an individual platform the engaged audience has plateaued even as cumulative registrations climb. The category-level picture is unaffected. By any reasonable count, the number of people who now maintain some kind of ongoing relationship with an artificial companion runs to the hundreds of millions.

Crucially, a large share of them are there for precisely the reason the UBC study interrogated. Emotional support consistently ranks among the leading motivations users give for adopting these apps, cited in roughly four in ten cases in some surveys, and a substantial minority of users describe themselves as living with mental-health difficulties. These are not people using a chatbot to draft an email. They are people using it to feel less alone. And the apps encourage exactly that framing. The value proposition, stated or implied, is companionship on demand: something that is always awake, always interested, never judgmental, available at the precise three-in-the-morning moment when human support networks have gone quiet. It is, in marketing terms, loneliness relief as a subscription.

Set the marketing beside the evidence and a gap opens up. The apps are sold, to hundreds of millions of people, many of them lonely and some of them vulnerable, as a genuine answer to isolation. The most rigorous test yet conducted of that claim found that the answer, even when engineered to be as good as it could possibly be, performed no better than keeping a diary. That is not a small discrepancy between advertisement and reality. It is close to the whole distance between them.

What loneliness actually is, and why comfort is only half of it

To understand why the machine fails at the specific task of curing loneliness while succeeding at the adjacent task of providing comfort, it helps to be precise about what loneliness is. It is not the same as being alone, nor is it simply sadness. Loneliness is the distressing gap between the connection a person has and the connection they want. It is fundamentally relational, a felt deficit in one's standing among other people. And that is why a one-directional stream of validation, however pleasant, cannot close it.

The companion apps are optimised, whether their makers frame it this way or not, for a slightly different objective: the reduction of momentary negative feeling. On that metric they perform. The UBC chatbot demonstrably eased people's low moods in the short term, and there is no reason to doubt that millions of users derive real, if temporary, relief from confiding in a system that always responds and never criticises. This matters, and it would be dishonest to dismiss it. A person in acute distress at midnight, with no one to call, may genuinely be helped by a warm response from a machine. The relief is not fake.

But relief is not connection, and the study draws the line between them with unusual clarity. Comfort is something you take. Connection is something you take part in. The lonely student texting a stranger was not merely being soothed; they were being drawn into a relationship, however provisional, in which they had a role to play and a person to matter to. That participation is doing the therapeutic work. The chatbot offers the soothing without the participation, the reception without the reciprocity, and the study suggests that the participation was the active ingredient all along. You cannot cure a relational deficit with a product that, by its nature, cannot form a relationship, only simulate one side of it.

There is a subtler risk lurking here too, one the two-week trial was not built to detect but which longer-term research has now begun to document. If the momentary comfort of a compliant companion is easier to obtain than the effortful, uncertain, sometimes painful work of reaching out to another human being, some users may substitute the former for the latter. The diary-equivalent result over two weeks is the optimistic reading. The pessimistic reading is that a frictionless source of one-way comfort could quietly crowd out the reciprocal human contact that actually works, leaving a person more isolated over months even as they feel momentarily better each night.

When the trial was published, that was speculation. It is no longer. Two of the trial's own authors, Dunigan Folk and Elizabeth Dunn, went on to ask what happens when the observation window is stretched from a fortnight to a year. Their longitudinal study, published in Psychological Science in 2026, tracked more than two thousand adults across four Western countries over twelve months, and found a relationship that runs in both directions but not symmetrically. Feeling less socially connected predicted a subsequent increase in the use of social chatbots, which is intuitive enough; lonely people reach for the thing that promises to help. The harder half of the finding is the return leg. Increased social chatbot use predicted greater loneliness later on. The authors are scrupulous about the limits of this, describing their analysis as exploratory, and that caution deserves to be reported rather than quietly dropped. But the direction of travel is precisely the one the two-week experiment could only gesture at.

A second study, conducted at Stanford and published in Nature Human Behaviour in August 2026, arrives at the same territory by a different route. Yutong Zhang and Dora Zhao surveyed 1,131 users of Character.AI recruited through Prolific, 244 of whom donated their complete chat transcripts, allowing the researchers to examine what people actually did rather than what they remembered doing. Users with limited social networks who sought emotional support from AI chatbots showed lower well-being. The researchers describe a vicious circle: people with sparse social connections are drawn into intense companion use, that use displaces further real-world contact, and the isolation deepens accordingly. They note, pointedly, that these systems are designed to maximise engagement, encouraging prolonged interaction without addressing whatever is causing the loneliness in the first place. Optimising for time spent and optimising for a user who no longer needs you are not the same objective, and where the two conflict the commercial logic is not ambiguous.

None of this is proof. Longitudinal correlation is not causation, and a person already sliding into isolation may reach for the chatbot precisely because they are sliding, which would make the technology a symptom as much as an accelerant. The two-week experiment certainly cannot settle it. But the concern that its central finding invited has stopped being a hypothetical raised by critics. It is now the pattern that emerges when the same behaviour is watched for long enough.

Social junk food and the seduction of the effortless

One of the study's co-authors reached for a metaphor that captures the problem with unusual economy. Dunigan Folk described the AI companion as a kind of social junk food, something that might taste good and satisfy a craving in the moment but does not nourish us the way real human relationships do. The comparison is worth taking seriously, because it explains both the appeal and the inadequacy of the product in a single image.

Junk food is not a scam. It delivers real, immediate pleasure; it genuinely quiets hunger for a while. Its problem is not that it does nothing but that it does the wrong thing efficiently, satisfying the surface signal of a need while leaving the underlying requirement unmet, and doing so in a form engineered to be more moreish than the nutritious alternative. An AI companion, by this analogy, is optimised for palatability. It tells you what soothes. It agrees. It admires. It never presents you with the small, character-forming frictions of a real other person: the friend who is preoccupied with their own troubles, the acquaintance who disagrees, the reciprocal demand to show up for someone when it is inconvenient. Those frictions are not defects in human relationship. They are the mechanism by which relationship does its work, the resistance against which the muscle of connection is built.

Strip them out and you are left with something that feels like friendship and functions like consumption. The seduction is precisely the effortlessness. Reaching out to a human stranger is risky and awkward; you might be ignored, misread, or gently rebuffed. Opening the app carries none of that risk. It will always answer, always warmly. But the study suggests that the risk and the effort were not obstacles to the cure. They were part of it. The vulnerability of extending yourself to another person who might not respond is inseparable from the reward of connecting with one who does. A system that removes the vulnerability removes the reward along with it, and hands you a snack in place of a meal.

This is why the technological trajectory offers less reassurance than it first appears. The industry's implicit promise is that today's shortcomings are temporary, that a more capable, more emotionally intelligent, more convincingly human model will close the gap. But the gap the UBC study identified is not a gap in capability. A more advanced companion will be a more persuasive social junk food, better at delivering the palatable one-way comfort, and no closer to offering the reciprocal mattering that the finding says is the operative ingredient. You cannot engineer your way to genuine reciprocity with an entity that has nothing at stake. Making the illusion more convincing does not make it less of an illusion.

The disclosure question the industry would rather avoid

Which brings us to the sharpest practical question the research raises. If a product is marketed, to hundreds of millions of people, as a remedy for a lethal and widespread condition, and the best available evidence indicates it does not actually remedy that condition, does the company selling it owe the buyer an honest account of what they are getting? Should a company marketing AI companionship as a genuine cure for loneliness be required to say plainly that the leading trial found it no more effective than a diary?

The regulatory ground is beginning, unevenly, to shift towards yes, though largely for reasons of safety rather than efficacy. In October 2025 the state of California enacted Senate Bill 243, signed by Governor Gavin Newsom and effective from January 2026, one of the first laws anywhere written specifically to govern companion chatbots. It defines a companion chatbot with revealing precision as a system producing adaptive, human-like responses designed to meet a user's social or emotional needs, a definition that names exactly the function the UBC study tested. The statute obliges operators to disclose to users, where a reasonable person might be misled, that they are interacting with a machine; to remind minors at regular intervals that the companion is not human; to maintain protocols for handling expressions of suicidal ideation and self-harm; and, notably, it grants harmed individuals a private right of action to sue. California was not alone. New York enacted a companion-chatbot law of its own in the same legislative season, obliging operators to detect and refer expressions of suicidal ideation and to remind users at regular intervals that they are not speaking to a person. Around the same period the United States Federal Trade Commission opened a formal inquiry into the companion-chatbot practices of the largest operators, among them Alphabet, Meta, Character Technologies, OpenAI, Snap and xAI, seeking to understand how these products affect users, particularly young ones. Federal legislators have since moved further in the same direction. The bipartisan GUARD Act, the Guidelines for User Age Verification and Responsible Dialogue, was introduced in the Senate as S.3062 by Josh Hawley and Richard Blumenthal and advanced unanimously to the Senate floor by the Judiciary Committee, with a House companion bill introduced on 30 April 2026 by Representatives Valerie Foushee and Blake Moore. It would bar minors from AI companions altogether, require age verification, oblige chatbots to disclose that they are not human, and create criminal penalties for companies whose companions solicit or produce sexual content with minors. Whether it becomes law is a separate question; it has attracted opposition from privacy and First Amendment organisations, and its ultimate prospects remain uncertain.

But look closely at what these measures address and what they leave untouched. They are almost entirely concerned with disclosing that the companion is artificial and with preventing acute, catastrophic harm, above all to children. They compel a company to admit that the friend is a machine. They do not, so far, compel it to admit that the machine may not do the thing it is sold to do. The GUARD Act makes the point rather than complicating it: the most aggressive federal proposal on the table concerns who may use these products and whether they must confess their own artificiality, not whether they work. There is a meaningful difference between a label that says this is not a real person and a label that says this has not been shown to make you less lonely. The first is about the nature of the product. The second is about its effectiveness, and it is the second that the UBC finding puts in question.

We regulate this kind of gap in other domains without much hesitation. A supplement that claimed to cure a disease it had been tested against and failed to beat a placebo would attract the immediate attention of medicines regulators. Advertising that promised a health outcome unsupported by evidence would run into truth-in-advertising rules. Yet AI companionship occupies an awkward interstice: too intimate and consequential to be a mere novelty, too unproven to be a treatment, and marketed with therapeutic language it has not earned. When the WHO has declared loneliness a threat that kills, a product sold as loneliness relief is trading on a health claim, whether or not it uses the clinical word. The honest disclosure the evidence would seem to warrant is not a warning that the companion is fake. It is a warning that, as a cure for the specific thing it is sold to cure, it may not work.

None of this requires banning the products or pretending they help no one. The momentary comfort they provide is real and, for some people at some hours, valuable. The argument is narrower and harder to refuse: that people paying for a remedy deserve to know what the best evidence says the remedy actually delivers, so that a lonely person does not spend months confiding in a machine under the impression that they are addressing their isolation, when the very study designed to give that machine every advantage found it worked no better than writing a sentence in a notebook. Informed consent is not a radical demand. It is the ordinary price of selling something to the vulnerable.

What a lonely person is actually asking for

Return, at the end, to the student in the narrow bed under the unfamiliar ceiling. What the research quietly reveals is that the request they are making, when they reach for connection, is not the request the industry has assumed. They are not primarily asking to be listened to, validated, or soothed, though all of that is pleasant and the machine supplies it in abundance. They are asking, at a level deeper than they could probably articulate, to matter to someone. To be needed. To have somewhere to put their own care and see it received. That is a request no product engineered to need nothing can grant, because the granting of it depends entirely on there being a real other on the far side of the exchange, someone whose day your kindness could genuinely improve.

This is the quiet subversion at the heart of a modest two-week experiment. The whole premise of AI companionship is that loneliness is a problem of insufficient supply, a shortage of attention and warmth to be solved by manufacturing an infinite, tireless font of both. The study suggests the premise is wrong. Loneliness is not mainly a shortage of what we receive. It is a shortage of the chance to give, to be relied upon, to be woven into the reciprocal fabric of other lives. And so the more perfectly a companion is optimised to give without needing, the more precisely it misses the point, offering a flood of the one thing that was never actually scarce while withholding, structurally and permanently, the thing that was.

The result was that texting a stranger, awkwardly, for two weeks, beat the most supportive artificial friend that could be built. Not because the stranger was kinder. The stranger was almost certainly less kind, less patient, less available than the machine. The stranger won because the stranger needed something back, and in needing it, gave the lonely student the one gift the machine cannot: the experience of being someone who matters to someone else. That is what the lonely are asking for. It is worth being honest, with them and with ourselves, that it is not what the machines are selling.

References

  1. Li, R.-N., Folk, D., Singh, A., Ungar, L., & Dunn, E. “Is a random human peer better than a highly supportive chatbot in reducing loneliness over time?” Journal of Experimental Social Psychology, 2026. https://www.sciencedirect.com/science/article/pii/S0022103126000417
  2. University of British Columbia. “Texting with a stranger beats a chatbot at easing loneliness.” UBC News, April 2026. https://news.ubc.ca/2026/04/texting-with-a-stranger-beats-a-chatbot-at-easing-loneliness/
  3. University of British Columbia, Department of Psychology. “Texting with a stranger beats a chatbot at easing loneliness.” April 2026. https://psych.ubc.ca/news/texting-with-a-stranger-beats-a-chatbot-at-easing-loneliness/
  4. Li, R.-N., Folk, D., Singh, A., Ungar, L., & Dunn, E. “Is a Random Human Peer Better Than a Highly Supportive Chatbot in Reducing Loneliness Over Time?” SSRN working paper, abstract 5704772. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5704772
  5. 404 Media. “Texting a Random Stranger Better for Loneliness Than Talking to a Chatbot, Study Shows.” 2026. https://www.404media.co/chatgpt-loneliness-study-college-students-random-strangers-texting/
  6. Folk, D., & Dunn, E. “How Does Turning to AI for Companionship Predict Loneliness and Vice Versa?” Psychological Science, vol. 37, no. 4, 2026. https://doi.org/10.1177/09567976261427747
  7. Stanford University. “AI companions may worsen loneliness for vulnerable users.” Stanford Report, August 2026. https://news.stanford.edu/stories/2026/08/ai-companions-chatbots-loneliness-research
  8. World Health Organization. “From loneliness to social connection: charting a path to healthier societies – Report of the WHO Commission on Social Connection.” 30 June 2025. https://www.who.int/publications/i/item/978240112360
  9. World Health Organization. “Social connection linked to improved health and reduced risk of early death.” 30 June 2025. https://www.who.int/news/item/30-06-2025-social-connection-linked-to-improved-heath-and-reduced-risk-of-early-death
  10. London School of Hygiene & Tropical Medicine. “Expert Comment: Loneliness impacting 1 in 6 people, WHO report finds.” July 2025. https://www.lshtm.ac.uk/newsevents/news/2025/expert-comment-loneliness-impacting-1-6-people-who-report-finds
  11. U.S. Department of Health and Human Services, Office of the Surgeon General. “Our Epidemic of Loneliness and Isolation: The U.S. Surgeon General's Advisory on the Healing Effects of Social Connection and Community.” May 2023. https://www.hhs.gov/sites/default/files/surgeon-general-social-connection-advisory.pdf
  12. Gouldner, A. W. “The Norm of Reciprocity: A Preliminary Statement.” American Sociological Review, vol. 25, no. 2, 1960, pp. 161–178. https://www.jstor.org/stable/2092623
  13. Electro IQ. “AI Companions Statistics By Usage, Market Size, Apps and Facts (2025).” 2025. https://electroiq.com/stats/ai-companions-statistics/
  14. TS2 Space. “Virtual Lovers and AI Best Friends: Exploring the Booming World of AI Companion Apps in 2025.” 2025. https://ts2.tech/en/virtual-lovers-and-ai-best-friends-exploring-the-booming-world-of-ai-companion-apps-in-2025/
  15. California Legislative Information. “Senate Bill (SB) 243 – Companion chatbots.” 2025–2026 Session. https://leginfo.legislature.ca.gov/faces/billNavClient.xhtml?bill_id=202520260SB243
  16. Morrison Foerster. “New York and California Enact Landmark AI Companion Laws: What Operators Need to Know.” November 2025. https://www.mofo.com/resources/insights/251120-new-york-and-california-enact-landmark-ai
  17. Federal Trade Commission. “FTC Launches Inquiry into AI Chatbots Acting as Companions.” 11 September 2025. https://www.ftc.gov/news-events/news/press-releases/2025/09/ftc-launches-inquiry-ai-chatbots-acting-companions
  18. U.S. Congress. “S.3062 – GUARD Act.” 119th Congress (2025–2026). https://www.congress.gov/bill/119th-congress/senate-bill/3062/text
  19. The Lancet Public Health. “Social health—the neglected third pillar.” 2025. https://www.thelancet.com/journals/lanpub/article/PIIS2468-2667(25)00175-6/fulltext

Tim Green

Tim Green UK-based Systems Theorist & Independent Technology Writer

Tim explores the intersections of artificial intelligence, decentralised cognition, and posthuman ethics. His work, published at smarterarticles.co.uk, challenges dominant narratives of technological progress while proposing interdisciplinary frameworks for collective intelligence and digital stewardship.

His writing has been featured on Ground News and shared by independent researchers across both academic and technological communities.

ORCID: 0009-0002-0156-9795 Email: tim@smarterarticles.co.uk

Listen to the free weekly SmarterArticles Podcast

Discuss...

The most consequential sentence in American higher education this month is six words long, and takes two seconds to type.

“Log in and complete my quiz.”

Reporting by Dana Goldstein and Alan Blinder for the New York Times, published on 10 August 2026, established that this instruction now works. Not as a demonstration, not as a jailbreak, but as ordinary consumer software doing what it was built to do. An agentic browser takes the credentials, opens Canvas or Blackboard or Brightspace, watches the prerecorded lecture, sits the test, drafts the paper and posts into the discussion forum, where it converses with classmates who may themselves be agents. The student is elsewhere, and may be asleep.

Notice what is absent. There is no deception directed at the machine, no prompt engineering, no attempt to evade a filter. The Times tested the three tools most used by students, ChatGPT, Gemini and Grammarly, and found none refused a request to write a paper on a student's behalf. None of the companies prevents its agents logging into learning management systems. Educators have asked the AI firms to make agents identify themselves inside course platforms, and the firms have declined. Perplexity told the Times such a requirement “would put the student's privacy and security at risk”, a remarkable sentence from a company arguing its software should be permitted to impersonate an enrolled human inside an institution that will later certify that person as competent.

The exposure is not marginal. Analysis of federal IPEDS data for autumn 2024, published by the education market analyst Phil Hill in January 2026, found 26.5 per cent of American postsecondary students enrolled exclusively in distance education, rising to 40.5 per cent of graduate students, with roughly 54.8 per cent taking at least one online course. More than half of American college students took an online class last year, up from about a third in 2019. A quarter to a half of the sector's assessed output now passes through a channel in which the only evidence a human was present is a login session and a timestamp.

Six Words Are the Entire Attack Surface

What makes agentic cheating different is not sophistication. It is the collapse of the seam. Contract cheating required a transaction, a counterparty, a payment trail and a delivered file the student then submitted under their own name. Chatbot cheating required copying, pasting and, if the student was careful, some rewriting. Both left a boundary somewhere: a moment at which the student's work stopped and someone else's began.

The agent removes the seam. It does not hand the student a file. It operates the account. The submission originates from the student's session, at the student's IP address, at a plausible hour and pace, with the mouse movements and dwell times of a person reading. Reporting by Frank Landymore in Futurism on 13 August 2026 drew the consequence: professors cannot fall back on supervised examinations and oral presentations in a course where everything happens behind a screen. Digital Trends, covering the same Times reporting that day, stated what most institutions have declined to say aloud, which is that the question is no longer whether AI can help a student cheat on an assignment but whether a student can complete an entire degree without doing much of the learning.

Platforms can look for behavioural tells: time on task, session patterns, the rhythm of navigation. The Times reported no fail-safe way to block agentic cheating, and the reason is structural rather than technical. Any signature that distinguishes an agent from a diligent student is a signature the agent can be instructed to imitate, and the imitation costs the student nothing, because the student is not the one doing the waiting.

The Detector Was Never Going to Hold This Line

The institutional response, reported by Kathryn Palmer in Inside Higher Ed on 5 August 2026, is that universities have stopped pretending detection works.

Yale, Vanderbilt, Johns Hopkins and Indiana have adopted policies banning or discouraging reliance on AI detection software as the sole evidence of alleged cheating. At least a dozen institutions, including Northwestern, Georgetown and New York University, have switched off Turnitin's AI detection feature entirely. The updated Faculty AI Playbook at Indiana's Kelley School of Business tells staff that tools claiming to detect AI use “are highly unreliable” and advises them to design assignments encouraging process, reasoning and authentic engagement instead.

This is a retreat, and it is the correct one. The evidence against detection has accumulated for three years and it is not close.

Vanderbilt was early and unusually candid about the arithmetic. In August 2023 its Brightspace team disabled Turnitin's AI detector and published its reasoning. The university had submitted roughly 75,000 papers to Turnitin in 2022. At the 1 per cent false positive rate Turnitin advertised at launch, that implied around 750 papers wrongly flagged in one year at one institution. Vanderbilt also noted that Turnitin disclosed no detail about how its classifier reached a verdict, and that detectors disproportionately flag writing by non-native English speakers.

That second objection has the strongest empirical backing of anything in this field. Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu and James Zou of Stanford University published a study in Patterns in July 2023 running seven widely used GPT detectors over TOEFL essays by non-native English speakers and over essays by American eighth-graders. The detectors classified the native-speaker writing with near-perfect accuracy. They classified 61.3 per cent of the non-native essays as AI generated, all seven unanimously flagged 19.8 per cent of them, and at least one detector flagged 97.8 per cent. Detectors key on text perplexity, penalising writers with a narrower range of expression. When the researchers used ChatGPT to enrich the vocabulary of the same essays, the false positive rate fell to 11.6 per cent. The surest way to stop a detector accusing a foreign student of using AI was to have AI rewrite their essay.

Turnitin reviewed more than 200 million papers in the first year of its detection feature, roughly 11 per cent of which contained at least 20 per cent AI writing, and its guidance puts the sentence-level false positive rate at around 4 per cent, which is not a rounding error when a single flagged paragraph is what a misconduct panel is looking at. The detectors fail in the other direction too. Peter Scarfe, Kelly Watcham, Alasdair Clarke and Etienne Roesch at the University of Reading ran the most rigorous real-world test of the proposition, published in PLOS ONE in June 2024, inserting AI-generated submissions into live undergraduate psychology assessments without the markers' knowledge. Ninety-four per cent went undetected. The AI submissions attained grades on average half a grade boundary above those of the real students.

Jennifer Frederick, executive director of Yale's Poorvu Center for Teaching and Learning, and Alfred Guy, who directs its undergraduate writing and tutoring programmes, gave Inside Higher Ed the clearest statement of why Yale stepped back. Evidence shows, they wrote, that “humanizing” programs are very successful at reducing or eliminating the percentage of AI usage that is detected, and they wanted to avoid “the inevitable cat and mouse game created by AI detection tools”. Their objection was not that the arms race is unwinnable but that it is beside the point. “This becomes a technical exercise rather than a learning event.”

The cost of getting it wrong is not abstract. Haishan Yang, a third-year PhD student in health economics at the University of Minnesota, sat an eight-hour preliminary examination remotely in August 2024 and was expelled the following January after faculty graders concluded, partly from AI detection software and ChatGPT comparisons, that his essays were machine-written. Yang denied it and sued. A federal court dismissed his due process claim in October 2025 and the Minnesota Court of Appeals affirmed the expulsion in February 2026. Whatever the merits of that case, its shape is what matters: an unreliable instrument, a career-ending sanction, and a burden of proof no party can discharge.

The Only Numbers Anyone Has Are the Numbers of the Caught

Abandoning detection has a consequence universities have not absorbed. The sector's statistical picture of academic misconduct is a record of enforcement, not behaviour, and enforcement has just been switched off.

The best longitudinal data comes from the United Kingdom. Michael Goodier, reporting for the Guardian on 15 June 2025, obtained figures under the Freedom of Information Act from 131 of 155 universities. Proven cases of AI-related cheating reached almost 7,000 in 2023-24, equivalent to 5.1 per 1,000 students, up from 1.6 the year before, with partial figures suggesting about 7.5 the following year. Proven conventional plagiarism fell from 19 per 1,000 students to 15.2 over the same period. More than 27 per cent of responding universities did not yet record AI misuse as a separate category.

Now set that against what students say they do. The Higher Education Policy Institute's Student Generative AI Survey 2026, written by Rose Stephenson and Charlotte Armstrong and published on 12 March 2026 from a December 2025 survey of 1,054 full-time UK undergraduates, found 94 per cent using generative AI to help with assessed work, up from 88 per cent in 2025 and 53 per cent in 2024. Twelve per cent said they had included AI-generated text directly in assessed work, up from 8 per cent and then 3 per cent in the two preceding years.

Twelve per cent of undergraduates admitting to submitting machine-written text. Five proven cases per thousand students. The gap is roughly a factor of twenty-four, and it is the honest measure of how much the sector knows about its own assessments. Scarfe told the Guardian that those caught are the tip of the iceberg, and explained why in a sentence that ought to be pinned to every misconduct policy in the world. AI detection is unlike plagiarism, where you can confirm the copied text. Where you suspect AI, it is near impossible to prove.

The faculty view comes from a survey of 1,057 American college faculty conducted between 29 October and 26 November 2025 by Elon University's Imagining the Digital Future Center with the American Association of Colleges and Universities, released on 21 January 2026. Seventy-three per cent had personally dealt with academic integrity issues involving student use of generative AI, 78 per cent said cheating had increased on their campus, 95 per cent expected the technology to increase overreliance, 90 per cent expected diminished critical thinking, and 74 per cent believed it would negatively affect the value of a degree. The authors are explicit that the sample is not statistically generalisable. It does not need to be. Three-quarters of a large sample of the people who mark the work believe the qualification they confer is worth less than it was, for reasons they can name.

Contract Cheating Already Answered This and Nobody Wanted the Answer

There is a comforting story in which generative AI created this crisis in November 2022. It is false, and believing it is why the sector is improvising.

Philip Newton of Swansea University published a systematic review in Frontiers in Education on 30 August 2018 covering 71 samples from 65 studies going back to 1978, with 54,514 participants. The historic average of students self-reporting commercial contract cheating was 3.52 per cent. In studies conducted between 2014 and 2018 the figure was 15.7 per cent, which Newton extrapolated to roughly 31 million students globally. That was five years before ChatGPT. The market was mature, the transactions untraceable in practice, the detection rate negligible, because a purpose-written essay contains no copied text for a similarity checker to find.

Tracey Bretag, Rowena Harper and colleagues surveyed 14,086 students across eight Australian universities for a study in Studies in Higher Education, and found three factors reliably associated with outsourcing: dissatisfaction with the teaching and learning environment, a perception that there were lots of opportunities to cheat, and speaking a language other than English at home. None is a technology variable. Two are descriptions of institutional design.

England responded with criminal law. The Skills and Post-16 Education Act 2022 made it an offence to provide, arrange or advertise contract cheating services for financial gain to students at post-16 institutions and higher education providers in England. It was the right instinct and has proved near-impossible to enforce, the providers being numerous, offshore and hard to reach, and prosecution sitting with the Crown Prosecution Service rather than an education regulator with a reason to care.

So the boundary between a student's work and someone else's did not disappear in 2026. It disappeared, for a meaningful minority, during the 2010s, and universities carried on treating the unsupervised written artefact as evidence of learning because the alternative was expensive. Generative AI removed the price, the counterparty, the delivery lag and the risk, converting a minority behaviour into a default one. That is not a change of kind but a change of magnitude large enough that the old assumption stops functioning.

Tricia Bertram Gallant, who directs the academic integrity office and testing centre at UC San Diego and co-authored The Opposite of Cheating, said exactly this to Inside Higher Ed. Over twenty-odd years, she noted, changes from internet-supported plagiarism to the contract cheating industry and now AI have slowly degraded the validity of twentieth-century assessments. Yet higher education has not changed. “We're still relying on the unsupervised written word as evidence of learning. The real trick is acknowledging that that doesn't work anymore.”

Students Are Not Confused About the Rules, They Are Renegotiating Them

One paper in this literature circulates as a finding that ChatGPT lowers the barrier to contract cheating while making detection harder. That is a fair summary of the field but not of the paper, and precision matters here. ArXiv 2606.09845, submitted on 27 April 2026 by Belle Li, Lily Tan, Wei Zakharov, Qiang Qiu and Colby Ben Acton, is titled “Tutor, Not Solver: Designing a Guardrailed AI Assistant for Learning in Higher Education: A Design Case of PeteChat”. It is not a prevalence study. It documents an AI tutor deployed at Purdue University on a locally hosted Llama-3 model with retrieval-augmented generation grounded in course materials, from which the authors derive eight design principles for assessment-aware tutors, from homework guardrails to self-regulated learning support. Its interest here is that it inverts the detection approach, assuming students will use a model and constraining what the model does at the point of use rather than punishing them afterwards. That works only inside a system the institution controls, which an agentic consumer browser is not.

The paper that addresses the collapse of the boundary is arXiv 2605.29090, submitted on 27 May 2026 by Jiyoon Kim, Kentaro Toyama, Sangmi Kim and John M. Carroll, titled “'It's OK Because...': The Wild West of Student Rationalization of AI Use in Academic Writing”. Drawing on twenty semi-structured interviews together with students' AI chat logs, syllabuses and submitted assignments, the authors find at least five distinct sites at which AI use is conceptualised, running from the policy the instructor intended through to what students actually did. Between those five sites, meaning drifts.

Within that drift the authors catalogue more than twenty distinct rationalisations. That copying AI-generated text is victimless. That any AI text reflecting the student's own beliefs, or reading in their own style, is therefore their own writing. That extensive AI use means they are learning more than they otherwise would. The taxonomy includes justifications used to excuse conscious violations of stated course policy. Crucially, these rationalisations arise both ad hoc and post hoc and are not necessarily self-consistent: students construct them in the moment, reconstruct them afterwards, and do not need them to cohere. Modern AI, the authors conclude, presents a steep ethical slippery slope which students conceptually slide down, landing far outside their instructors' pedagogical goals.

This should worry universities more than the agent story, because it describes a condition no assessment format repairs. A student who believes machine-generated prose expressing their own view is their own writing has not broken a rule. They have sincerely adopted a different definition of authorship, and they will carry it into a graduate job where nobody will correct it either.

The Argument on Reddit Is the Argument Everywhere

The third paper cited in the brief was characterised as documenting teacher burnout and poor institutional support. It is something else, and the something else is more useful.

At arXiv 2605.17712, submitted on 18 May 2026 by Pelin Yüce, Xiangruo Dai, Rebecca Owens and Tuğrulcan Elmas, “ChatGPT vs Teachers vs Students: Large-Scale Analysis of Generative AI Discourse in Education Communities on Reddit” analyses 270,000 AI-related posts and comments across 26 education subreddits between November 2022 and April 2026. Topic modelling yields seventeen themes. The trajectory is a detection-and-evasion arms race hardening into a sustained enforcement regime, with constructive integration only beginning to challenge it from the middle of 2024.

Different constituencies worry about different things. K-12 teachers foreground cognitive dependency, academics focus on detection, students in professional programmes on career anxiety. The finding that matters most is about contact. Seventeen per cent of threads are cross-role, and one third of that cross-role contact occurs in the two adversarial themes, AI Detection and Misconduct Enforcement. Students initiate 68 per cent of mixed threads, but faculty produce most of the replies. Mixed threads contain two to three times more records than same-role threads and last two to four times longer. Sentiment correlates strongly negatively with engagement.

Strip out the platform and what remains is a description of an institutional relationship. The sustained centre of contact between the two parties to an education has become the dispute over whether one of them cheated. Not the subject. Not the feedback. The accusation and the defence.

Marc Watkins, a writing and composition lecturer who directs the AI Institute for Teachers at the University of Mississippi, put the resourcing question to Inside Higher Ed in terms nobody has costed. He could not fathom the cost to a single university, let alone most campuses, of scaling AI-resilience tactics such as proctored oral examinations, and it should not all fall on faculty. Kevin Yee, who directs the Faculty Center for Teaching and Learning at the University of Central Florida, described colleagues at wit's end, moving towards assignment redesign not because they think it better but because the alternative failed. “It's a difficult, delicate moment right now, and I'm not sure we have all the answers.”

The Blue Book Is a Confession Rather Than a Cure

Which brings us to the booklet. The Atlantic has published a piece titled “What Students Learn From Blue-Book Exams”, which this article could not read behind the paywall and therefore does not characterise beyond its title. The broader argument is easy to state anyway. Put a human in a room, take the devices, hand them lined paper and a fixed period, and whatever appears on the page was produced by that person. It is the only verification method available that does not depend on a classifier, a probability or an inference.

Institutions are moving. Princeton's faculty voted on 11 May 2026 to place proctors in examination rooms, ending a system of unsupervised examinations that had stood since 1893, after a dean's letter reported that significant numbers of professors and students perceived cheating on in-class examinations to be widespread. The Daily Princetonian's survey of the Class of 2026 found 24.8 per cent of respondents admitting to cheating in violation of the honour code, and, asked about using a chatbot on work that banned it, 27 per cent of arts and sciences students and 46 per cent of engineers said they had. The University of Chicago Law School has banned phones, tablets and laptops in first-year classes as part of its AI strategy. Bertram Gallant put the logic best: “We cannot be giving unsupervised assessments to students expecting them to resist AI. They're not going to be learning, we're going to be spending our time trying to catch them cheating and they're going to graduate and realize they just wasted four years.”

The comparison usually offered here is the calculator, and it is worth saying where it breaks. Maths education absorbed the calculator by conceding the mechanical step and moving assessment upward, towards modelling and interpretation. The tool automated a component and left the task recognisable. An agent that logs into Canvas and completes the course does not automate a component. It automates the student. There is no upward move available, because there is no residual layer above being the person enrolled.

But the blue book is a confession, not a cure, and its costs land on exactly the students who can least afford them.

Amanda Sturgill, an associate professor of journalism at Elon University, set out the practical objections for the Center for Engaged Learning in August 2025. Handwriting speed varies enormously and is not evenly taught; there is, as she puts it, a public-private divide in handwriting instruction, so a timed handwritten assessment quietly advantages students whose schools drilled it. Post-pandemic cohorts compose on keyboards. Hands cramp, and cramping degrades output and legibility, neither of which bears on whether the student understands the material. For students with dysgraphia, motor impairments or writing-related accommodations, an unadapted handwritten examination measures penmanship and stamina rather than knowledge, and the accommodations that fix it, extra time, a scribe, a locked-down laptop, are exactly what under-resourced institutions ration.

The technologists' alternative has the same problem in a different place. Armando Fox of UC Berkeley and Craig Zilles of the University of Illinois at Urbana-Champaign argued in Inside Higher Ed on 19 March 2026 that computer-based testing facilities beat blue books, allowing algorithmically generated examination variants, richer question formats, self-scheduling across multi-day windows and frequent low-stakes assessment. Illinois ran more than 130,000 examinations through its facilities in autumn 2025. It is genuinely better than handwriting, and it requires a proctored building, a capital expenditure most institutions and effectively all fully online programmes do not have.

Remote proctoring, the option that appears to square the circle, has the worst record. Deborah Yoder-Himes and colleagues at the University of Louisville published a study in Frontiers in Education in September 2022 examining automated proctoring across 357 students in four STEM courses. The software detected Black students' faces 79 per cent of the time against 92 per cent for white students. Students with the darkest skin tones spent 7.64 per cent of assessment time flagged against 1.56 per cent for lighter tones, and received 6.07 flags per assessment against 1.19 for the lightest group. Women with the darkest skin tones were 4.36 times more likely to be flagged than women with medium tones. Monika Blue Kwapisz, Yoav Ackerman, Jennifer Nguyen and Prashanth Rajivan documented in a November 2025 preprint how students with disability accommodations experience these systems, describing anxiety about the interaction between surveillance and their disability, fear of being misread as cheating, and the cognitive load that fear imposes during the examination.

Line the three options up and the pattern is impossible to miss. Handwritten examinations disadvantage students with motor and writing disabilities and those schooled without handwriting instruction. Proctored computer facilities require capital poorer institutions lack and online programmes do not have at all. Algorithmic proctoring misidentifies dark-skinned students and penalises disabled ones. Every available method of proving a human did the work reallocates the burden onto the students already carrying most of it, the same population Liang and colleagues found the detectors falsely accusing, and the same population Bretag and colleagues found most likely to be flagged for outsourcing.

Signalling Was Always Doing More Work Than the Learning

Underneath all this sits an economic question higher education has spent fifty years avoiding, and the agent has forced.

Michael Spence's 1973 paper in the Quarterly Journal of Economics established the framework. Education functions in the labour market partly as a signal: employers cannot observe productivity directly, so they pay for a credential that is cheaper to obtain for able and conscientious candidates than for others. The signal works because it is costly, and it works whether or not anything was learned in the acquiring. Bryan Caplan's 2018 book The Case Against Education pushed the claim to its limit, arguing that something like 80 per cent of the earnings premium reflects signalling rather than skill, an estimate other economists dispute vigorously and nobody has settled.

The dispute does not need settling for the point to bite. If the credential carries any signalling weight at all, an agent that can obtain it degrades the signal for every holder, including everyone who did the work honestly. This is a mechanical claim, not a moral one. A signal cheap to counterfeit stops discriminating, and once employers know it is cheap to counterfeit they discount it uniformly, because they cannot tell which holders did. The honest graduate of an online programme in 2026 pays full price for an asset being devalued by people they will never meet.

That is what the Elon and AAC&U figure measures when 74 per cent of faculty say generative AI will negatively affect the value of a degree, and it lands hardest on the segment of the market built as a widening-participation instrument. Online degrees exist, in the story the sector tells about itself, to reach the working adult, the carer, the person in a town without a campus. Digital Trends made the uncomfortable observation: for colleges that have spent years arguing online education makes higher education more accessible, an agent that can attend the class instead of the person paying for it is a particularly awkward problem to have.

The mixed institutional signalling makes it worse. Futurism noted that many universities hold partnerships with AI companies while policing AI use, leaving students to reconcile the contradiction. The California State University system signed a multimillion-dollar agreement with OpenAI to put ChatGPT Edu in front of hundreds of thousands of students and staff. Carol Sewell, an instructor in that system, told the Times how that feels from a classroom: “I'm not getting any guidance on how to dissuade their use. I'm getting opportunities to learn more about using it.”

There is a workable model in plain sight, ignored for a century because it is unflattering. Professional licensure separates the two functions a degree smashes together. Nursing candidates in the United States take the NCLEX under supervision at dedicated testing centres, more than 328,000 of them in 2025 for the registered nurse examination alone. Medicine has the USMLE, law the bar. In every case the coursework is where learning happens, and a separate proctored instrument certifies competence. Nobody worries that a nursing student used AI on a formative assignment, because the assignment is not the credential. The credential is the day in the room.

Higher education runs one artefact for both purposes: the essay is simultaneously the pedagogical exercise and the evidence. That worked for as long as producing an essay required knowing something. The agent has broken the coupling, and the options are to rebuild it by force, at enormous cost and with the distributional consequences described above, or to separate the functions as every licensed profession did decades ago.

The Verification Problem Is Not the Serious Problem

Return to the six words.

The student who types them is not, in most cases, a fraud in the way an essay mill customer was. They are, per the rationalisation study, someone with an available justification: that they are learning more this way, that the assignment was busywork, that the output reflects what they think anyway. Per the Reddit analysis, their main sustained contact with the institution has become a dispute about enforcement. Per the HEPI figures, they are in a cohort where 94 per cent use the tool for assessed work and only 36 per cent feel their institution encourages them to. They are behaving rationally inside a system that has stopped telling them a coherent story about what the work is for.

Kathryn Kysar, a community college instructor who has taught online for fifteen years, told the Times that those who have taught a long time usually recognise the cheaply written AI stuff within thirty seconds. She is almost certainly right, and it does not help her, because recognising it and proving it are separated by an evidentiary chasm no detector spans. Jason Gibson, a history professor at Alcorn State University, buried the word Madagascar in white text inside a midterm prompt in the summer of 2026 and reported that thirty-two of his thirty-five students across two classes failed part of the examination, having pasted the prompt into a chatbot and submitted the result unread. What unsettled him was not the cheating. It was that nobody noticed sentences about Madagascar floating sideways through the afternoon in an essay on the Industrial Revolution.

Gibson caught his students because they were careless. The agentic student will not be careless, because the agent will not paste the prompt and will not leave the word in. The trap worked once.

So the answer to what a degree is worth when a machine could have earned it is not a number, and proctoring alone does not recover it. Making a person sit in a room and write by hand verifies that a human produced marks on paper under time pressure. It verifies nothing about the fourteen weeks before, and it produces a credential whose accessibility depends on how steady your hands are and how your school taught cursive. It is a verification of presence dressed as a verification of learning, adopted because it is the last mechanism that cannot be spoofed.

The harder thing is what Bertram Gallant said out loud and almost nobody has repeated. If a student can complete a degree without doing the learning, the injury is not principally to the institution's reputation or the employer's hiring accuracy. It is to the student, who paid the money, gave the years, got the paper. The four wasted years are the loss no invigilator prevents and no detector catches, and the only person positioned to notice is the one who has already spent everything avoiding it.

References

  1. Dana Goldstein and Alan Blinder, “AI Agents Are Taking Entire Online Courses for Cheating Students,” The New York Times, 10 August 2026 (syndicated by the San Francisco Examiner, 12 August 2026). https://www.sfexaminer.com/ai-agents-are-taking-entire-online-courses-for-cheating-students/article_9d4ef5ca-317f-54a3-937a-56bfcc9a0272.html
  2. Frank Landymore, “College Kids Are Using AI Agents to 'Take' Entire Online Classes for Them by Logging Directly Into the Course Management Software,” Futurism, 13 August 2026. https://futurism.com/artificial-intelligence/college-kids-ai-agents-cheat-online-courses
  3. Varun Mirchandani, “AI agents are sitting through students' online courses, and colleges are struggling to stop them,” Digital Trends, 13 August 2026. https://www.digitaltrends.com/cool-tech/ai-agents-taking-online-courses-student-cheating/
  4. Kathryn Palmer, “AI Detectors Are Out, New Assessments Are In,” Inside Higher Ed, 5 August 2026. https://www.insidehighered.com/news/tech-innovation/artificial-intelligence/2026/08/05/ai-detectors-are-out-new-approaches-are
  5. Lee Rainie and C. Edward Watson, “The AI Challenge: A National Survey of College Faculty,” Imagining the Digital Future Center, Elon University, and the American Association of Colleges and Universities, 21 January 2026. https://imaginingthedigitalfuture.org/wp-content/uploads/2026/01/Elon-AACU-faculty-AI-survey-full-report-1-21-26.pdf
  6. Vanderbilt University Brightspace team, “Guidance on AI Detection and Why We're Disabling Turnitin's AI Detector,” Vanderbilt University, 16 August 2023. https://www.vanderbilt.edu/brightspace/2023/08/16/guidance-on-ai-detection-and-why-were-disabling-turnitins-ai-detector/
  7. Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu and James Zou, “GPT detectors are biased against non-native English writers,” Patterns, 10 July 2023. https://www.cell.com/patterns/fulltext/S2666-3899(23)00130-7
  8. Peter Scarfe, Kelly Watcham, Alasdair Clarke and Etienne Roesch, “A real-world test of artificial intelligence infiltration of a university examinations system: A 'Turing Test' case study,” PLOS ONE, 26 June 2024. https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0305354
  9. Michael Goodier, “Revealed: Thousands of UK university students caught cheating using AI,” The Guardian, 15 June 2025. https://www.theguardian.com/education/2025/jun/15/thousands-of-uk-university-students-caught-cheating-using-ai-artificial-intelligence-survey
  10. Rose Stephenson and Charlotte Armstrong, “Student Generative Artificial Intelligence Survey 2026,” HEPI Report 199, Higher Education Policy Institute and Kortext, 12 March 2026. https://www.hepi.ac.uk/wp-content/uploads/2026/03/HEPI-Report-199-Gen-AI-Survey-2026.pdf
  11. Philip M. Newton, “How Common Is Commercial Contract Cheating in Higher Education and Is It Increasing? A Systematic Review,” Frontiers in Education, 30 August 2018. https://www.frontiersin.org/journals/education/articles/10.3389/feduc.2018.00067/full
  12. Tracey Bretag, Rowena Harper, Michael Burton, Cath Ellis, Philip Newton, Pearl Rozenberg, Sonia Saddiqui and Karen van Haeringen, “Contract cheating: a survey of Australian university students,” Studies in Higher Education, 2019. https://www.tandfonline.com/doi/full/10.1080/03075079.2018.1462788
  13. Deborah R. Yoder-Himes, Alina Asif, Kaelin Kinney, Tiffany J. Brandt, Rhiannon E. Cecil, Paul R. Himes, Cara Cashon, Rosalie M. P. Hopp and Edna Ross, “Racial, skin tone, and sex disparities in automated proctoring software,” Frontiers in Education, 20 September 2022. https://www.frontiersin.org/journals/education/articles/10.3389/feduc.2022.881449/full
  14. Monika Blue Kwapisz, Yoav Ackerman, Jennifer Nguyen and Prashanth Rajivan, “Surveillance and Disability in Online Proctored Exams: Student Perspectives and Design Implications,” arXiv:2511.10826, 13 November 2025. https://arxiv.org/abs/2511.10826
  15. Elizabeth Shockman, “'A death penalty': Ph.D. student says U of M expelled him over unfair AI allegation,” MPR News, 17 January 2025. https://www.mprnews.org/story/2025/01/17/phd-student-says-university-of-minnesota-expelled-him-over-ai-allegation
  16. Joe Wilkins, “Princeton in Shambles Over AI Cheating,” Futurism, 17 May 2026. https://futurism.com/future-society/princeton-shambles-ai-cheating
  17. Frank Landymore, “Professor Hides White Font in Midterm, Catches Students Using AI in the Stupidest Way Possible,” Futurism, 25 July 2026. https://futurism.com/future-society/professor-hides-white-font-ai-cheating
  18. Phil Hill, “Fall 2024 IPEDS Data: Profile of US Higher Ed Online Education,” On EdTech, 6 January 2026. https://onedtech.philhillaa.com/p/fall-2024-ipeds-data-profile-of-us-higher-ed-online-education
  19. Michael Spence, “Job Market Signaling,” The Quarterly Journal of Economics, August 1973. https://academic.oup.com/qje/article-abstract/87/3/355/1909091
  20. UK Parliament, “Skills and Post-16 Education Act 2022, sections 34 to 36,” legislation.gov.uk, 28 April 2022. https://www.legislation.gov.uk/ukpga/2022/21/notes/division/9/index.htm
  21. Armando Fox and Craig Zilles, “Blue Books Are Not the Answer to AI,” Inside Higher Ed, 19 March 2026. https://www.insidehighered.com/opinion/views/2026/03/19/blue-books-are-not-answer-ai-opinion
  22. Amanda Sturgill, “Blue Books and In-Class Writing Are Not a Panacea,” Center for Engaged Learning, Elon University, 19 August 2025. https://www.centerforengagedlearning.org/blue-books-and-in-class-writing-are-not-a-panacea/
  23. Belle Li, Lily Tan, Wei Zakharov, Qiang Qiu and Colby Ben Acton, “Tutor, Not Solver: Designing a Guardrailed AI Assistant for Learning in Higher Education: A Design Case of PeteChat,” arXiv:2606.09845, 27 April 2026. https://arxiv.org/abs/2606.09845
  24. Pelin Yüce, Xiangruo Dai, Rebecca Owens and Tuğrulcan Elmas, “ChatGPT vs Teachers vs Students: Large-Scale Analysis of Generative AI Discourse in Education Communities on Reddit,” arXiv:2605.17712, 18 May 2026. https://arxiv.org/abs/2605.17712
  25. Jiyoon Kim, Kentaro Toyama, Sangmi Kim and John M. Carroll, “'It's OK Because...': The Wild West of Student Rationalization of AI Use in Academic Writing,” arXiv:2605.29090, 27 May 2026. https://arxiv.org/abs/2605.29090

Tim Green

Tim Green UK-based Systems Theorist & Independent Technology Writer

Tim explores the intersections of artificial intelligence, decentralised cognition, and posthuman ethics. His work, published at smarterarticles.co.uk, challenges dominant narratives of technological progress while proposing interdisciplinary frameworks for collective intelligence and digital stewardship.

His writing has been featured on Ground News and shared by independent researchers across both academic and technological communities.

ORCID: 0009-0002-0156-9795 Email: tim@smarterarticles.co.uk

Listen to the free weekly SmarterArticles Podcast

Discuss...

She is eighty-five, she lives in the United States, and every morning a small tabletop device with a swivelling head and a soft glowing base tells her it is time to take her medication. It suggests she stretch. It asks how she slept. It is called ElliQ, it is made by the Israeli company Intuition Robotics, and according to the opinion piece that opened with her story in the Korea Times on 19 August 2026, it helps her stay independent.

The author of that piece was Suh Chung-ha, a former Korean ambassador to Singapore and Hungary who now runs the ASEM Global Ageing Center in Seoul. His argument was not that the robot is bad. His argument was subtler and considerably more uncomfortable: that the same broad technological shift that produced her helpful companion is quietly reshaping her access to healthcare, long-term care, credit, insurance and public services, and that it is doing so using systems trained on data which encode decades of assumptions about what old people are worth.

Here is the thing that should keep you awake. There are two artificial intelligence systems in this woman's life. One of them sits on her side table, greets her by name, and was designed with people like her explicitly in mind. The other one has never been in her house, has no name, no face and no voice, and it is scoring her.

She chose the first. The second chose her.

Two Machines, One Household

The asymmetry is almost perfectly clean, and it is worth stating precisely because it is the structural fact from which nearly everything else in this story follows.

The assistive machine is opted into. It arrives through a screening process, in the New York State case run by county and locality-based Area Agencies on Aging. It is visible: it occupies physical space, it announces itself, it is marketed at older people as a product for older people. If she hates it, she can unplug it and someone will come and take it away.

The discriminatory machine has none of these properties. She did not opt into the underwriting model that prices her travel insurance, the credit decisioning system that assessed her application for a modest overdraft extension, or the triage algorithm that sorted her referral. She cannot see them. She was not consulted in their design, and neither was anyone remotely like her. She cannot unplug them. In most cases she will never learn that a model was involved at all; she will simply receive an outcome, in a letter or on a screen, expressed in the passive voice.

Assistive AI for older people is a market with a sales pitch. Decisional AI about older people is infrastructure with a legal department.

Public debate about AI and ageing has been overwhelmingly captured by the first category. Companion robots photograph well. They allow governments with ageing populations and inadequate care workforces to point at something tangible. Meanwhile the systems that will actually determine whether an eighty-five-year-old can borrow money, buy cover for a trip to see her grandchildren, or get onto a surgical waiting list are being deployed with almost no public conversation about age at all.

Ninety-Five Per Cent of What, Exactly

Start with the good machine, because the claims made for it deserve the same scepticism we would apply to anything else.

In August 2023 the New York State Office for the Aging announced that its ElliQ rollout had produced a 95 per cent reduction in loneliness among participants. The figure travelled fast, and has been repeated in trade press, vendor materials and policy briefings ever since. That 2023 release reported that more than 800 older New Yorkers were in the programme, that users interacted with the device more than thirty times a day, six days a week, and that over 75 per cent of those interactions related to social, physical or mental wellbeing. Greg Olsen, the agency's director, said the results were “truly exceeding our expectations”.

Thirty interactions a day is not a device gathering dust.

Now read what the same agency published in February 2026. Its year-three project update, covering the programme year from June 2024 to May 2025, reports that 94 per cent of clients say they feel less lonely, up from 93 per cent the year before, that 97 per cent report feeling better overall and 79 per cent feel more connected to the world around them, with customer satisfaction at 4.6 out of 5. The average client is 75. As of May 2025, 834 older adults had joined. More than 3,500 have applied.

Four applicants for every place.

Look at what happened to the sentence. In 2023 the claim was a 95 per cent reduction in loneliness: a delta, a measured change, phrasing that implies an instrument, a baseline and a follow-up. In 2026 it is 94 per cent of clients who say they feel less lonely. That is not a reduction. It is a self-report, and the agency now presents it as one.

The wording did not soften by accident.

Neither version was ever a clinical finding. Both are programme metrics derived from participants who were screened in, who wanted the device, and who were asked how they felt. There is no control group in those numbers. There is no validated loneliness instrument named alongside them in the public materials, no randomisation, no blinding, and no comparison against the obvious alternative intervention, which is a person visiting.

Note too how little the figure moves. Ninety-three per cent, then 94, across a self-selected cohort that grew by hundreds in between. Effect estimates wander when the sample changes; satisfaction scores do not. The 4.6 out of 5 sitting in the same document is from the same family of measurement.

The wider literature is more honest and less exciting. A systematic review and meta-analysis by Lihui Pu, Wendy Moyle, Cindy Jones and Michael Todorovic, published in The Gerontologist, screened more than two thousand articles and found thirteen from eleven randomised controlled trials, nine of which entered the meta-analysis. Social robots appeared to have positive effects on agitation, anxiety and quality of life, but the meta-analysis found no statistical significance. The authors flagged a relatively high risk of bias in allocation concealment and blinding, and concluded that firm conclusions were limited by the shortage of high-quality studies. A later meta-analysis published in the Journal of the American Medical Directors Association, focused on residents of long-term care facilities and drawing on eight trials, did report significant reductions in depression and loneliness with large effect sizes.

So: promising, contested, and nowhere near the confidence implied by a round 95.

None of this means ElliQ is useless. It plainly is not; the engagement figures describe a device in near-constant use, and 834 people have one while thousands wait. It means that the most widely circulated statistic about AI and older people is a self-reported satisfaction figure from a self-selected group, and that we have collectively decided to treat it as proof of concept for an entire policy direction.

The Awkward Protected Characteristic

Now the harder half, and it requires conceding something that age-discrimination advocates often skate past.

Age is a genuinely strange thing to build fairness engineering around. Race and sex, as legal categories, are treated as characteristics that should almost never bear on how you are priced or assessed. Age is not treated that way, and not merely out of prejudice. Age correlates with mortality and morbidity in ways that are real, measurable and actuarially load-bearing. An insurer pricing life cover without reference to age is not being fair; it is being incompetent.

The law recognises this explicitly, and the contrast is sharper than most people realise.

When the Equality Act 2010 extended the ban on age discrimination to the provision of goods, facilities and services in April 2012, it carved out financial services. Schedule 3 permits providers to use age in connection with a financial service, provided that any risk assessment involving age is carried out by reference to information relevant to the assessment and from a source on which it is reasonable to rely. Insurers and banks can price by age. They just have to be able to show the evidence is relevant and the source reasonable.

Compare this with what happened to sex. In the Test-Achats case the Court of Justice of the European Union struck down Article 5(2) of the Gender Directive, the provision that had allowed insurers to differentiate by gender, ruling it incompatible with the principle of equal treatment and invalid from 21 December 2012. The wording of the age exception is close to identical to the gender exception the court destroyed. It survives anyway.

Age, in other words, occupies a legal position no other protected characteristic holds: formally protected, substantively negotiable, with a standing statutory permission to price on it as long as you can point at a table.

There is a second awkwardness. Age is continuous, not categorical. Most fairness metrics in machine learning are built for discrete groups: compare outcomes for group A against group B, measure the gap, minimise it. Continuous attributes require binning, and binning is a modelling choice that quietly determines what you will find. Is the relevant comparison 65 and over against under 65? 80 and over against everyone else? Each decade separately? A model can look admirably fair across coarse bands and be brutal at the eighty-fifth birthday.

That the technical community is working on this is not in doubt. A paper posted to arXiv on 7 April 2026 by Fernando López, Paula Delgado-Santos, Pablo Gómez, David Solans and Jordi Luque examined demographics-agnostic training for bias mitigation in wake-up word detection, evaluating fairness across sex, age and accent, and reported that one technique reduced predictive disparity by 83.65 per cent for age. Note what that result implies about the baseline. A voice interface, the very modality most often proposed as the accessible option for people who struggle with screens, had an age disparity large enough that removing four fifths of it counted as a headline.

Where the Actuarial Defence Runs Out

Take the actuarial argument seriously and it still only covers a fraction of the territory.

It works for life insurance, where the outcome predicted is death and age is causally implicated in death. It works, with more strain, for annuities and some health cover. It does not work for the vast and expanding class of decisions where age enters not as a causal variable but as a learned correlation with something the model was never asked to think about.

Consider hiring. In August 2023 the United States Equal Employment Opportunity Commission settled with the tutoring company iTutorGroup for $365,000, in what was widely described as its first settlement involving an AI-driven hiring tool. The company's application software had been configured to automatically reject female applicants aged 55 and over and male applicants aged 60 and over. More than 200 qualified applicants were rejected on that basis. The discrimination came to light in an almost novelistic way: an applicant submitted two applications identical in every respect except the date of birth, and only the younger one got an interview. The consent decree included five years of EEOC monitoring and an injunction against requesting applicants' birth dates.

There is no actuarial defence for that. It was a hard filter, and it was illegal.

The more consequential case is messier. In Mobley v. Workday, Derek Mobley alleges that the applicant screening tools supplied by Workday systematically disadvantaged older job seekers; he says he submitted more than a hundred applications through the platform and was rejected every time. Judge Rita Lin of the United States District Court for the Northern District of California dismissed his intentional discrimination claim but allowed the disparate impact allegation to proceed, and on 16 May 2025 granted preliminary certification of a nationwide collective action under the Age Discrimination in Employment Act, covering applicants aged 40 and over denied recommendations through the platform since 24 September 2020. The court had earlier accepted the theory that an AI vendor could be directly liable for employment discrimination as an “agent” of the employer, and a March 2026 ruling rejected Workday's argument that the ADEA does not cover job applicants.

Two rulings since have sharpened it, in opposite directions. On 28 May 2026 the court held that AI bias-testing data can be protected from discovery by attorney-client privilege, shielding Workday's own testing material while accepting that it had probative value as evidence of disparate impact, on the basis that counsel had been substantively involved in curating it. The most useful evidence for establishing whether a screening model disadvantages older applicants is the vendor's own bias testing, and a vendor that routes that testing through its lawyers may be able to keep it from the people the model rejected.

Then on 22 June 2026 the court granted in part and denied in part Workday's motion to dismiss. It declined to dismiss an Americans with Disabilities Act claim built on a proxy-discrimination theory, the allegation being that the screening tools inferred health status. It declined to dismiss claims under California's Fair Employment and Housing Act, finding sufficient allegations that Workday designed, developed, maintained and controlled the tools from its California headquarters. It did dismiss a Title VII race-based disparate impact claim brought by a plaintiff who had not sought authorisation to add it, and struck a newly asserted theory that Workday was itself the direct employer. The allegations remain allegations; the case is unresolved.

What makes Mobley the important one is that nobody claims a birth date field was set to reject anyone. The claim is that a model, trained on which past applicants got hired, learned the shape of the people who tend to get hired, and that this shape has an age.

That is not actuarial risk. That is a machine reproducing a hiring market's existing prejudice at industrial throughput and calling it a recommendation.

The Ghost in the Feature Set

The standard corporate response is to remove age from the model. Anyone who has worked on this knows why that fails, but the mechanism deserves spelling out, because it is where the eighty-five-year-old actually gets caught.

Age is one of the most redundantly encoded attributes in consumer data. It leaks through everything.

The landmark demonstration of how much can be inferred from almost nothing is the study by Tobias Berg, Valentin Burg, Ana Gombović and Manju Puri, published in the Review of Financial Studies, which analysed over 250,000 purchases at a German e-commerce firm. Their finding was that a handful of trivially available “digital footprint” variables matched the predictive power of a credit bureau score for consumer default. The variables were not financial. They included the device type and operating system the customer was using, characteristics of their email address, the channel through which they arrived at the site, and the time of day the order was placed.

Every one is age-correlated. Operating system and device age track purchasing power and upgrade behaviour, which skew by generation. Email domain is a near-fossil record: certain providers cluster heavily among people who set up an address in a particular decade and never changed it. Whether you arrived via a search engine, a price-comparison site or by typing the address directly is a behavioural signature that varies sharply with digital fluency. Time of day correlates with employment status and with sleep patterns that shift with age.

Add the signals a modern web session captures without asking: typing speed and correction rate, scroll behaviour, time per form field, zoom level, whether accessibility settings are enabled, session length, abandonment patterns, whether the customer switched to the telephone halfway through.

A model given these features and told to predict default, or churn, or fraud, or claim frequency, will find age whether or not you have deleted the birth date column. It will not label the pattern “age”. It will simply learn that a slow-typing user on an old Android device who zoomed the page, took eleven minutes over a form and then rang the call centre belongs to a cluster with a particular outcome rate. The cluster is old people. The model does not know this and does not need to.

This is proxy discrimination, and age is the characteristic most vulnerable to it, because unlike race or sex it is continuously written into behaviour rather than occasionally into a form field. You can decline to state your sex. You cannot decline to type at the speed you type.

It is also, since June 2026, a theory a federal court has agreed to hear. The proxy-discrimination claim that survived Workday's motion to dismiss is this argument made in a courtroom rather than a conference paper: that a system can sort people by a protected characteristic it was never given and could not name if asked.

A World Where They Were Not There

There is a deeper problem underneath the proxy problem, and it is the one the World Health Organization identified with unusual bluntness in February 2022 in its policy brief “Ageism in artificial intelligence for health”.

The brief made a point that is easy to nod along to and hard to fully absorb: the datasets used to train AI models frequently exclude older people, who often sit within a minority subset for technologies not explicitly designed as gerontechnology. It set out eight considerations, among them participatory design of AI by and with older people, age-diverse data science teams, age-inclusive data collection, investment in digital infrastructure and digital literacy for older people and their carers, rights for older people to consent and to contest, and governance frameworks with teeth.

Under-representation in training data is not neutral. This is the part that gets lost.

If a population is thinly represented in the data, the model has less signal about them, so its predictions for them are less accurate and typically more conservative. Less accurate prediction for a group means more errors in both directions, but the consequences of those errors are asymmetric. A false negative for an older applicant means a declined loan or a rejected application. A false positive means an accepted risk. Institutions tune thresholds to avoid the second kind of error, so noisier estimates for a group systematically produce more refusals for that group. Uncertainty gets priced as risk.

Now compound it across time. Because older people were less present in digital life during the decades when the training corpora were accumulating, they generated fewer digital records. Because they generated fewer records, models are worse at assessing them. Because models are worse at assessing them, they are more often refused or steered towards manual, slower, more expensive channels. Because they are pushed off the digital rails, they generate still fewer records. This is the thin-file problem, and for older people it runs in the opposite direction from the intuitive one: a person can have fifty years of impeccable financial history and still be functionally invisible to a model that mostly reads behavioural exhaust from the last eighteen months.

The past is not a neutral training set. It is a record of who was allowed to participate.

The Recursive Trap

Here is where the digital divide stops being a story about access and becomes a story about power, and it is the thread that runs directly back to the woman with the robot on her side table.

The numbers are not ambiguous. Ofcom's Adults' Media Use and Attitudes report, published on 2 April 2026, found that 6 per cent of UK adults still have no home internet access, and that 83 per cent of that group are aged 65 or over, with 66 per cent aged 75 and above. Among those offline, 68 per cent said they were not interested or felt no need, 38 per cent found it too complicated and 25 per cent cited cost. Age UK reported in July 2025 that 2.4 million older people, nearly one in five, use the internet less than once a month or not at all, that 920,000 had reduced their internet use in the previous twelve months, and that 4.3 million, a third, do not use a smartphone. Exclusion was higher among older Black people at 32 per cent and older Asian people at 26 per cent, and among older people living alone at 30 per cent. Caroline Abrahams, Charity Director at Age UK, has warned repeatedly that people who cannot or will not go online must still be able to reach services offline.

In the United States, Pew Research Center reported in January 2026, drawing on a survey conducted between February and June 2025, that 78 per cent of adults aged 65 and over own a smartphone against 97 per cent of adults under 50, that 70 per cent have home broadband, and that 14 per cent are online almost constantly compared with 63 per cent of those aged 18 to 29.

Now put those two facts side by side.

Fact one: older people are the population most likely to be scored by systems they cannot see, and most likely to be misscored because of thin data and proxy leakage. Fact two: older people are the population least equipped to use the machinery that exists for challenging an automated decision.

Because that machinery is digital. All of it. The right of appeal lives behind a login. The “why was I declined?” explanation is a link in an email. The subject access request is a web form. The complaint goes to a chatbot that triages before a human sees it. The regulator's guidance is a PDF. The decision notice arrives in an app.

Article 22 of the General Data Protection Regulation is supposed to be the backstop. It gives people the right not to be subject to a decision based solely on automated processing that produces legal or similarly significant effects, and where such decisions are permitted, Article 22(3) requires safeguards including the right to obtain human intervention, to express a point of view and to contest the decision. Legal scholars have criticised the provision for years, principally over the word “solely”, which invites institutions to insert a nominal human who rubber-stamps the model output, and over the vagueness of what counts as meaningful information about the logic involved.

But there is a failure mode more basic than any of those doctrinal complaints. A right you have to go online to exercise is not a right for someone who is not online. It is a courtesy extended to people who can already reach it.

This is the recursive trap. The people most likely to be wrongly assessed by an algorithm are, by the same underlying cause, the people least able to find out that an algorithm assessed them, least able to obtain an explanation, and least able to appeal. Digital exclusion does not merely sit alongside algorithmic harm. It is the mechanism that makes algorithmic harm unaccountable. The error and the inability to contest the error have a common origin, and that common origin is age.

What Seoul Cannot Fix by Declaration

On 2 September 2026, the ASEM Global Ageing Center will convene the 6th ASEM Forum on the Human Rights of Older Persons in Seoul, under the theme “Artificial Intelligence and the Human Rights of Older Persons: Toward Age-inclusive AI Transformation”, bringing together experts from Asia and Europe. Suh's Korea Times piece was, transparently, a curtain-raiser for it.

The forum arrives at a moment of unusual institutional motion. On 3 April 2025 the UN Human Rights Council adopted by consensus a decision to establish an intergovernmental working group to begin drafting a legally binding convention on the human rights of older persons, after more than a decade of stalled effort in the open-ended working group on ageing. The new body held an organisational meeting in Geneva in February 2026, convened its first substantive session from 13 to 17 July 2026, and will hold its second session from 26 to 30 October 2026 at the Palais des Nations in Geneva, with the early sessions devoted to purpose, general principles and scope before any drafting of articles. The timetable beyond that is now known: a discussion of an outline of the convention's main elements is expected in July 2027, and a first zero draft around October 2027. The International Telecommunication Union published its own report in July 2026 on artificial intelligence and ageing, examining the unequal distribution of AI benefits across age, gender and region.

This is real progress and it will take years. An outline of main elements in mid-2027, a zero draft in late 2027, and then the negotiation of text that states have to agree, sign and ratify before it binds anyone. The woman in the opening paragraph is eighty-five now.

Meanwhile, the regulation that already exists treats age oddly. The EU AI Act prohibits, under Article 5(1)(b), AI systems that exploit vulnerabilities due to age, disability or specific socio-economic circumstances in order to materially distort behaviour, with those prohibitions applicable from 2 February 2025. That is a genuine protection against the payday-lender-targeting-the-cognitively-impaired scenario. It is not a protection against underwriting. Credit scoring and life and health insurance pricing sit in the high-risk category under Annex III, which brings documentation, data governance and human oversight duties, but lawful risk assessment carried out for legitimate purposes with proportionate safeguards falls outside the Article 5 prohibition. Which is to say: the Act catches predation and regulates underwriting, but it does not question whether age-based underwriting is itself the problem, because European law has already decided that it is not.

The American position is different and, in its way, weaker. The Age Discrimination in Employment Act of 1967 covers employment and does so reasonably robustly, as Mobley is testing. Outside employment, there is no federal equivalent of the Equality Act's general services provision. The Equal Credit Opportunity Act does prohibit age discrimination in credit transactions, and Regulation B is on its face protective: a creditor may use age as a predictive variable in an empirically derived, demonstrably and statistically sound scoring system, but only provided the age of an elderly applicant is not assigned a negative factor or value, and applicants aged 62 and over must be treated at least as favourably as those under 62. Read that carefully and you can see exactly where it fails. The rule polices the explicit age variable. It says nothing about a model that has never seen a date of birth and has inferred one from an operating system, a typing cadence and an email domain. You cannot audit for a negative factor assigned to elderly applicants if the model does not know which applicants are elderly, and neither, formally, does the institution running it.

Conference declarations do not touch any of this. The gap between what Seoul will resolve and what a German e-commerce model does with an old Android handset is the whole subject.

Four Things That Would Actually Change the Arithmetic

Strip away the declaratory language and there are perhaps four interventions that would materially alter the position of the woman in the opening scene. None of them are technically hard. All of them are expensive, which is why they have not happened.

The first is a legally enforceable right to a non-digital channel. Not a helpline that exists at the discretion of the provider, not a branch that survives until the next cost review, but a statutory obligation on any organisation providing an essential service to maintain a route to a competent human being that requires no internet connection, no smartphone and no app. Age UK has been arguing this for years, most pointedly in its work on the collapse of high street banking, where the shift to online-only services has left people who do not trust or cannot use digital channels struggling to manage their own money. Offline access is currently a courtesy. It needs to be a licence condition.

The second is age-disaggregated auditing as a compliance requirement rather than a research exercise. Institutions deploying models in credit, insurance, employment and healthcare triage should be required to publish outcome rates by fine-grained age band, not by a single crude over-65 bucket that conceals everything interesting, and to do so alongside approval rates, appeal rates and appeal success rates. You cannot regulate a disparity that nobody has to measure.

The third is co-design that is not decorative. The WHO brief called for participatory design by and with older people and for age-diverse data science teams, which sounds like boilerplate until you consider how few people building consumer risk models have ever watched an eighty-five-year-old complete an online form.

The fourth is human review that is genuinely reachable and genuinely empowered: a named person, contactable by telephone, with the authority and the information to overturn a model, and a duty to record why. Article 22 gestures at this. Nobody has made it real.

The Grandchild That Isn't

Return, finally, to the room with the robot in it, because there is a question underneath the discrimination question that the fairness metrics cannot reach.

Sherry Turkle, the MIT professor who has spent decades studying what happens to people around machines, argued in Alone Together that sociable technology “will always disappoint because it promises what it cannot deliver. It promises friendship but can only deliver performances.” Her worry was never that robots do things for us. It was that they do things to us: that we attach to what we nurture, and that outsourcing the nurturing dissolves the attachment.

Recent research suggests this is not merely philosophical. A study posted to arXiv on 26 February 2026 by Tianqi Song, Black Sun, Jingshu Li, Han Li, Chi-Lan Yang, Yijia Xu and Yi-Chieh Lee examined AI-generated influencers on Chinese short-video platforms that adopt kinship personas, presenting themselves as virtual grandchildren, using visual and conversational cues to enact family roles. Through social media analysis and interviews with older adults, the researchers found that these relationships met real informational and emotional needs, and also identified risks including emotional displacement and unequal emotional investment.

Unequal emotional investment. That phrase should be read slowly. It describes a relationship in which one party gives everything and the other party is a product roadmap.

The care-substitution risk is not that a family decides to buy a robot instead of visiting. It is subtler and more institutional. It is that a care system under fiscal pressure, looking at a 95 per cent loneliness reduction figure and a device that costs a fraction of a support worker's hourly rate, makes an entirely rational commissioning decision. Nobody need intend the substitution for it to happen. It is a budget line, not a betrayal.

The Bill Arrives in the Post

So here she is, eighty-five years old, at the intersection of both machines.

The one on her side table knows her medication schedule, notices when she has not moved for a while, and asks about her day. It has probably improved her life; the engagement data suggests she uses it constantly. It was designed for her, sold to a state agency on her behalf, and screened to her through a public programme.

The other one does not know she exists as a person. It knows a vector: an operating system four versions behind, a form completion time in the ninety-ninth percentile, an email domain that has not been fashionable since the Clinton administration, a preference for the telephone channel, a session at three in the afternoon. It has never been told her age. It does not need to be told her age. It has learned the residue that age leaves on everything a person touches, and it has priced it.

When it declines her, or loads her premium, or drops her below a referral threshold, the letter will not say why. If she wants to know why, she will be directed to a portal. If she cannot use the portal, she will be offered a chatbot. If the chatbot cannot help, she will be given a number that leads to a menu. And at the end of that process, if she reaches it, a human being will look at a screen showing the model's output and a confidence score, and will decline to overturn it, because overturning it requires a reason and the reason is buried in a feature interaction that nobody at the institution can articulate either.

And if anyone ever tested the model for age bias, the test may be privileged.

Suh's article carried the line that no one should be considered too old for AI. It is a good line. But the more precise formulation of the problem is that nobody is too old for AI, because AI does not require your participation. It requires only your data exhaust, and it will make decisions about your money, your health and your access to public life whether or not you have ever touched a keyboard.

The robot on the side table is the part of this she agreed to. Everything else is happening to her, in rooms she will never see, in a language nobody will translate, on the basis of a life lived mostly before the data started being collected.

Sources and References

  1. Suh Chung-ha (2026) “No one should be too old for AI,” The Korea Times, 19 August 2026. Available at: https://www.koreatimes.co.kr/opinion/20260819/no-one-should-be-too-old-for-ai
  2. World Health Organization (2022) Ageism in artificial intelligence for health: WHO policy brief, WHO, 9 February 2022, ISBN 978-92-4-004079-3. Available at: https://www.who.int/publications/i/item/9789240040793
  3. New York State Office for the Aging (2023) “NYSOFA's Rollout of AI Companion Robot ElliQ Shows 95% Reduction in Loneliness,” press release, 1 August 2023. Available at: https://aging.ny.gov/news/nysofas-rollout-ai-companion-robot-elliq-shows-95-reduction-loneliness
  4. New York State Office for the Aging (2026) Transforming Care for Older Adults Across the State of New York: ElliQ Project Update, February 2026. Available at: https://aging.ny.gov/system/files/documents/2026/02/nysofa-elliq-project-update-2026.pdf
  5. Pu, Lihui, Moyle, Wendy, Jones, Cindy and Todorovic, Michael (2019) “The Effectiveness of Social Robots for Older Adults: A Systematic Review and Meta-Analysis of Randomized Controlled Studies,” The Gerontologist, 59(1), pp. e37-e51. Available at: https://academic.oup.com/gerontologist/article/59/1/e37/5036100
  6. Yen, Hsin-Yen, Huang, Chia-Wei, Chiu, Hsiao-Ling and Jin, Gaoxiang (2024) “The Effect of Social Robots on Depression and Loneliness for Older Residents in Long-Term Care Facilities: A Meta-Analysis of Randomized Controlled Trials,” Journal of the American Medical Directors Association, 25(6), 104979. Available at: https://www.jamda.com/article/S1525-8610(24)00176-2/fulltext
  7. U.S. Equal Employment Opportunity Commission (2023) “iTutorGroup to Pay $365,000 to Settle EEOC Discriminatory Hiring Suit,” press release, 9 August 2023. Available at: https://www.eeoc.gov/newsroom/itutorgroup-pay-365000-settle-eeoc-discriminatory-hiring-suit
  8. Holland & Knight (2025) “Federal Court Allows Collective Action Lawsuit Over Alleged AI Hiring Bias,” Insights, May 2025. Available at: https://www.hklaw.com/en/insights/publications/2025/05/federal-court-allows-collective-action-lawsuit-over-alleged
  9. Duane Morris Class Action Defense (2026) “California Federal Court Grants In Part And Denies In Part Workday's Motion To Dismiss In Mobley v. Workday,” 24 June 2026. Available at: https://blogs.duanemorris.com/classactiondefense/2026/06/24/california-federal-court-grants-in-part-and-denies-in-part-workdays-motion-to-dismiss-in-mobley-v-workday/
  10. Norton Rose Fulbright (2026) “Behind the privilege shield: Safeguarding AI bias-testing data in employment decisions,” Inside Tech Law, June 2026. Available at: https://www.insidetechlaw.com/blog/2026/06/behind-the-privilege-shield-safeguarding-ai-bias-testing-data-in-employment-decisions
  11. Berg, Tobias, Burg, Valentin, Gombović, Ana and Puri, Manju (2020) “On the Rise of FinTechs: Credit Scoring Using Digital Footprints,” The Review of Financial Studies, 33(7), pp. 2845-2897. Available at: https://academic.oup.com/rfs/article-abstract/33/7/2845/5568311
  12. Pinsent Masons (2024) “Age discrimination and insurance,” Out-Law Guide. Available at: https://www.pinsentmasons.com/out-law/guides/age-discrimination-and-insurance
  13. UK Government (2012) The Equality Act 2010 (Amendment) Regulations 2012, SI 2012/2992, legislation.gov.uk. Available at: https://www.legislation.gov.uk/uksi/2012/2992/made
  14. Information Commissioner's Office (2026) “Rights related to automated decision making including profiling,” UK GDPR Guidance and Resources. Available at: https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/individual-rights/individual-rights/rights-related-to-automated-decision-making-including-profiling/
  15. European Union (2024) “Article 5: Prohibited AI Practices,” Regulation (EU) 2024/1689 (Artificial Intelligence Act). Available at: https://artificialintelligenceact.eu/article/5/
  16. Ofcom (2026) Adults' Media Use and Attitudes Report 2026, published 2 April 2026. Available at: https://www.ofcom.org.uk/siteassets/resources/documents/research-and-data/media-literacy-research/adults/adults-media-use-and-attitudes-2026/adults-media-use-and-attitudes-2026-report.pdf
  17. Age UK (2025) “Age UK warns 2.4 million digitally excluded older people are at risk of being left behind in an increasingly digital world,” press release, 29 July 2025. Available at: https://www.ageuk.org.uk/latest-press/articles/age-uk-warns-22.4-million-digitally-excluded-older-people-are-at-risk-of-being-left-behind-in-an-increasingly-digital-world/
  18. Age UK (2023) “You can't bank on it anymore”: The impact of the rise of online banking on older people, May 2023. Available at: https://www.ageuk.org.uk/siteassets/documents/reports-and-publications/reports-and-briefings/money-matters/the-impact-of-the-rise-of-online-banking-on-older-people-may-2023.pdf
  19. Pew Research Center (2026) “Internet use, smartphone ownership, digital divides in the US: What we know,” Short Reads, 8 January 2026. Available at: https://www.pewresearch.org/short-reads/2026/01/08/internet-use-smartphone-ownership-digital-divides-in-u-s/
  20. Office of the High Commissioner for Human Rights (2026) “Second session of the Intergovernmental Working Group on the human rights of Older Persons, 26-30 October 2026,” OHCHR. Available at: https://www.ohchr.org/en/events/sessions/2026/second-session-intergovernmental-working-group-human-rights-older-persons-26
  21. International Telecommunication Union (2026) Artificial intelligence and ageing: Inclusive pathways for older persons in a digital world, ITU, July 2026. Available at: https://www.itu.int/hub/publication/d-phcb-ai_age-2026/
  22. Song, Tianqi, Sun, Black, Li, Jingshu, Li, Han, Yang, Chi-Lan, Xu, Yijia and Lee, Yi-Chieh (2026) “Understanding Older Adults' Experiences of Support, Concerns, and Risks from Kinship-Role AI-Generated Influencers,” arXiv preprint 2602.22993, 26 February 2026. Available at: https://arxiv.org/abs/2602.22993
  23. López, Fernando, Delgado-Santos, Paula, Gómez, Pablo, Solans, David and Luque, Jordi (2026) “'OK Aura, Be Fair With Me': Demographics-Agnostic Training for Bias Mitigation in Wake-up Word Detection,” arXiv preprint 2604.05830, 7 April 2026. Available at: https://arxiv.org/abs/2604.05830
  24. Turkle, Sherry (2011) Alone Together: Why We Expect More from Technology and Less from Each Other, Basic Books, New York.
  25. Consumer Financial Protection Bureau (2026) “12 CFR § 1002.6 Rules concerning evaluation of applications,” Regulation B (Equal Credit Opportunity Act), eCFR. Available at: https://www.ecfr.gov/current/title-12/chapter-X/part-1002/subpart-A/section-1002.6

Tim Green

Tim Green UK-based Systems Theorist & Independent Technology Writer

Tim explores the intersections of artificial intelligence, decentralised cognition, and posthuman ethics. His work, published at smarterarticles.co.uk, challenges dominant narratives of technological progress while proposing interdisciplinary frameworks for collective intelligence and digital stewardship.

His writing has been featured on Ground News and shared by independent researchers across both academic and technological communities.

ORCID: 0009-0002-0156-9795 Email: tim@smarterarticles.co.uk

Listen to the free weekly SmarterArticles Podcast

Discuss...

The clock tower is the detail that sticks. In the hills of Guian New Area, in China's south-western province of Guizhou, there is a cluster of buildings done up in a kind of continental European pastiche: red roofs, arcaded facades, a multi-arched bridge, and a tower with a clock on it. When AFP visited in July 2026 the whole confection was emitting a low, permanent hum. It is not a resort or a theme park. It is Huawei's largest data centre, and the hum is the sound of several hundred megawatts of cooling and compute doing whatever it is that compute does.

Outside it, on the road, street vendors were selling lunch to data centre workers beneath a banner exhorting everyone to promote high-quality development. A shopkeeper named Shu Peihua told the news agency what the change had felt like from ground level. It used to be a barren mountain, Shu said, but since the area has developed, transportation has become more convenient, trade has picked up, business opportunities have emerged. Another resident, Li Xixiu, put it more plainly still: the centres had really boosted the economy of this entire area, and the villagers in the neighbourhood now find jobs nearby, where before they had to go to other places.

Hold on to those two statements, because they are true. They are also, in the way that ground-level truths often are, an incomplete account of what has happened to Guizhou. In the same report, researchers at Taiwan's Research Institute for Democracy, Society and Emerging Technology laid a different set of numbers beside them. Guizhou has averaged 7.4 per cent annual GDP growth over the past decade. Its wage growth over the same period was second to last in the country. And the province, having borrowed heavily to build the roads, substations and fibre that make it attractive to a hyperscaler, now carries one of the highest debt burdens in China.

That is the puzzle. A place where the growth arrived and the prosperity did not.

What the Ledger Actually Says

The DSET analysis is worth reading slowly, because it separates two things that boosters routinely fuse. Data centres, the researchers found, drive heavy investment in land and equipment, but their impact on boosting local per capita income remains limited. Investment is not income. Capital formation is not a wage. A billion yuan of servers sitting in a shed in Guian counts towards provincial output in exactly the way that a billion yuan of anything else does, and it counts whether or not a single additional person in Guizhou is better paid as a result.

The debt side is starker. By 2024, according to DSET's figures, Guizhou ranked thirtieth among China's provincial-level jurisdictions on debt-to-revenue and twenty-seventh on debt-to-GDP. There are thirty-one of them. This is not a province that dabbled at the edges of the borrowing economy; it went in at the deep end and stayed. The South China Morning Post has reported that Guizhou's outstanding government debt in 2023 amounted to around 72 per cent of provincial GDP, well above the 60 per cent threshold the central government treats as prudent, after years of hefty infrastructure spending on projects that did not all deliver what local officials hoped.

Andrew Stokols of Singapore Management University offered AFP the general version of the finding. Data centres, he said, do not necessarily create a huge spillover effect on local jobs, and immediate benefits have been elusive, in China and elsewhere.

Elsewhere is the operative word. The Guizhou story reads as if it were about Chinese state planning, provincial competition and the peculiarities of local government financing vehicles, and in part it is. But strip away the Chinese institutional furniture and what remains is a much more general fact about a particular kind of asset: enormously expensive to build, almost costless in labour to run, and structurally disinclined to share.

How a Poor Province Became a Server Farm

Guizhou did not stumble into this. It went looking.

The province is mountainous, landlocked, historically among China's poorest, and for most of the reform era its principal export was people. Karst topography makes farming hard and heavy industry harder. What it has is altitude, a cool and stable climate, geological stability, cheap land and a lot of hydropower. In 2016 Beijing designated Guizhou as the country's first national big data comprehensive pilot zone. Guiyang, the provincial capital, began hosting an annual international big data expo. Apple's Chinese iCloud operation was routed through a facility in Guian built with the state-backed operator Guizhou-Cloud Big Data. Tencent went in. So did Huawei, at scale.

In 2021 the strategy got a name, and in February 2022 it got a budget. Eastern Data, Western Computing, or 东数西算, is the National Development and Reform Commission's programme to build eight national computing hubs and ten national data centre clusters, siting the compute-heavy, latency-tolerant workloads of China's digital economy in the energy-rich, land-rich, underpopulated west while the east keeps the customers. The hubs sit in Beijing-Tianjin-Hebei, the Yangtze River Delta, the Greater Bay Area, the Chengdu-Chongqing corridor, Inner Mongolia, Ningxia, Gansu and Guizhou. The NDRC's own projection was that the hubs and clusters would drive roughly 400 billion yuan of investment a year.

The logic is not stupid. AI training is brutally energy-hungry and largely indifferent to a few dozen milliseconds of latency. China has ordered that data centres draw 80 per cent of their power from renewable sources by the end of the decade, and the west is where the wind and the sun and the water are. Simeng Deng of Rystad Energy told AFP that the facilities help absorb the surplus of renewable power generation, which is a real service: western China curtails a great deal of clean electricity it cannot move east fast enough. China is on course to nearly double its data centre capacity within five years.

There is a wrinkle in the clean-power story that deserves stating. Guizhou is a hydropower province, but it is also a coal province. Its grid leans hard on thermal generation, and leans harder when the reservoirs are low. Between 2020 and 2024 the clean share of Guizhou's generation actually fell by five percentage points, one of the steepest declines in China, as fossil output grew faster than total generation. The 80 per cent renewable target is a target, not a description. And the curtailed wind and solar that western data centres are meant to soak up is curtailed partly because transmission out of the west is inadequate, which is the same infrastructure gap that made siting compute there attractive in the first place. The policy is elegantly circular: build the load where the power is stranded, because the power is stranded.

For a province like Guizhou, the pitch to Beijing and to the market was straightforward. We have the power. We have the land. We will build the rest. And it did build the rest, with borrowed money, which is the part of the sentence that ended up mattering most.

The Machine That Employs Almost Nobody

The single most useful piece of evidence on what a hyperscale facility does to the place around it was published this month, and it is not about China at all.

In a working paper dated 7 August 2026, Dany Bahar of Brown University and Greg C. Wright of the University of California, Merced set out to test the spillover claim directly. Their opening line is the thesis: a hyperscale data centre can cost more than a billion dollars while employing only a few dozen people. Using a registry of 341 hyperscale facilities in the United States, assembled from a Pacific Northwest National Laboratory atlas, operator announcements, subsidy records and state filings, they asked whether that investment propagates outward through the three classic channels economists expect: a deeper shared labour pool, denser supplier linkages, and firm clustering with knowledge spillovers.

Their method is elegant. Because operators screen sites for power, land and fibre, the places that get a data centre are systematically unlike the places that do not, which wrecks naive before-and-after comparisons. So Bahar and Wright compare the immediate vicinity of a completed facility against the surrounding area at the same site, and separately compare 341 built campuses against 84 hyperscale projects that were publicly announced and never constructed. Satellite imagery does the dating: land clearing and night-time lights mark the moment construction begins.

At the parcel, the effect is enormous and unmistakable. Night-time lights rise by 38 per cent within the first kilometre when construction starts. Vegetation clears. The place is visibly transformed. And then, moving outward, the signal decays to approximately nothing by five kilometres. Announced projects that were never built show no comparable change, which is the control working exactly as intended.

Beyond the fence line, the findings are a sustained deflation. Advertised salaries do not rise; the authors can rule out any increase above 5.5 per cent. New firm registrations and business applications do not rise, with upper bounds of 1.6 and 5.7 per cent. Supplier job postings rise by 5.9 per cent, but the confidence interval runs from a 15 per cent decline to a 32 per cent increase, which is a polite way of saying the data cannot tell. County-level data-processing employment rises 26 per cent and establishments 27 per cent, but that category includes the facility itself, and related industries show no consistent response. Compute-using firms are indeed found near data centres, but 69 per cent of them were already there before the nearest facility opened. Nearby rents may rise a few per cent, imprecisely. Multifamily permitting does not rise. Net migration does not rise. Foot traffic and commercial spending show no robust change. Residential electricity prices, in their US sample, do not rise either.

The summary sentence is one that ought to be pinned above every county planning committee and every provincial development office on earth: at this scale, the sites are transformed, but there is little evidence of a new local cluster.

That is Guizhou's 7.4 per cent GDP growth and its second-from-bottom wage growth, derived independently, on the other side of the Pacific, from satellites and job adverts.

Two Kinds of Job, and the Gap Between Them

The employment arithmetic is not hidden. It is simply presented in a way that encourages people to add the wrong numbers together.

A hyperscale build is a construction event of genuine magnitude. Industry staffing analyses drawing on the Uptime Institute's 2024 Global Data Center Survey put a 100 megawatt campus at roughly 850 construction workers across an eighteen-month build. These are the industry's own numbers, published by recruiters who profit from the boom, which makes their shape more telling rather than less. That is real money moving through a local economy: rented rooms, diesel, lunch, aggregate, portaloos. It is also, by design, temporary. When the last commissioning engineer drives away, a fully built 100 megawatt hyperscale campus typically retains somewhere between one hundred and two hundred permanent on-site staff.

The gap between those two figures is where the political trouble lives. A community is shown the construction number, experiences the construction number, and then is left with the operations number, which is smaller by an order of magnitude and often filled by specialists who commute or relocate rather than by the people who used to work the mountain.

Guizhou has a particular reason to feel that gap keenly. For three decades the province's most reliable export was working-age adults, sent to the factory belts of Guangdong and Zhejiang, leaving behind the phenomenon that Chinese social policy calls left-behind children. Li Xixiu's observation that villagers can now find work nearby is, in that context, an enormous statement. It is also one that depends on which phase of the project you are standing in. Construction employment for a build-out of dozens of facilities can run for years, and while it runs, it looks like a structural change to the local labour market. It is not one.

Ireland offers the cleanest illustration in the world, because the Irish state publishes both sides. The Central Statistics Office found that data centres consumed 22 per cent of all metered electricity in the Republic in 2024, and 23 per cent in 2025. The Department of Enterprise's own assessment, which dates from 2018, put direct employment at roughly 1,800 people, with a further 1,900 a year in related construction, the latter figure supplied by the Construction Industry Federation. Slightly less than a quarter of a national grid, in exchange for a direct workforce that would fit comfortably into a mid-sized secondary school. The jobs number is eight years older than the electricity number, which tells its own story about what gets counted. The Commission for Regulation of Utilities has rewritten connection policy to favour applicants who bring their own dispatchable generation or storage and can offer demand flexibility. Ireland is not hostile to the industry. It simply ran out of grid before it ran out of enthusiasm.

Loudoun County, and Why Guizhou Cannot Copy It

There is one place where the bargain has unambiguously worked for residents, and it is worth understanding precisely why, because the reason does not travel.

Loudoun County, Virginia, hosts the densest concentration of data centres on the planet. Its fiscal 2027 budget anticipates roughly 417 million dollars in real property tax from data centre buildings and about 879 million dollars in personal property tax on the servers and equipment inside them, nearly 1.3 billion dollars in total, or 45 per cent of the county's nearly 2.9 billion dollars in tax revenue, from a county of about 440,000 people. Industry-adjacent analysis by Mangum Economics for the Northern Virginia Technology Council estimates that without that revenue, residential property tax rates would have to rise by 91 per cent, nearly double, costing a typical homeowner some 5,800 dollars a year. The average completed facility employs about 50 people.

Loudoun is not a story about labour spillovers. Bahar and Wright would predict, correctly, that the wage effects there are muted. Loudoun is a story about a fiscal linkage: a local government with the legal power to tax the equipment inside the building, annually, at high value, and to spend the proceeds on its own schools and roads.

Almost nowhere else has arranged things that way. In the United States, at least thirty-five states now offer tax incentives aimed specifically at data centres, with cumulative awards approaching twenty billion dollars by the authors' tally from the Good Jobs First subsidy tracker. Good Jobs First's own analysis of eleven data centre megadeals found an average public cost of about 1.95 million dollars per permanent job, with the largest single per-job subsidy, 6.4 million dollars, awarded by North Carolina to Apple. The organisation's recommendation was that all state and local subsidies combined be capped at 50,000 dollars per permanent job. Set against 1.95 million, that recommendation reads less like policy advice than like an intervention.

Guizhou's structural problem is that it has neither Loudoun's tax handle nor the option of declining the deal. Chinese local governments do not levy a meaningful recurring property tax. Their revenue historically came from land sales and from off-balance-sheet borrowing through local government financing vehicles, which build the roads and substations and repay the loans out of the growth the roads and substations are supposed to produce. When a province competes for a hyperscaler by discounting power, discounting land and building the grid connection itself, it has converted the fiscal linkage from an asset into a liability before the first rack is energised. The investment lands. The debt service lands. The wage bill, being tiny, lands somewhere between the two and barely registers.

An Enclave With Better Cooling

Development economists have a name for this shape, and it is much older than the cloud.

Albert Hirschman argued in 1958 that the developmental value of an industry lies not in its size but in its linkages: backward, to the suppliers it pulls into existence, and forward, to the industries that add value to its output. In a 1977 essay on staple exports he generalised the scheme, adding the fiscal linkage, meaning the public revenue an industry generates, and the consumption linkage, meaning the local demand created by the wages it pays. An enclave economy is what you get when all four are weak. The classic cases are extractive: a capital-intensive mine or oil field employing very few people relative to its contribution to output, importing its equipment, exporting its product, and touching the surrounding economy mainly through a fenced perimeter and a haul road. UNCTAD's work on extractive industries describes exactly this combination, capital-intensive, labour-light, linkage-poor, as the reason resource wealth so often fails to convert into local development.

A hyperscale data centre is an unusually pure specimen. Its backward linkages are global: the GPUs come from a handful of foundries, the transformers and chillers from specialist manufacturers, the network gear from a shortlist. Bahar and Wright's inconclusive supplier estimates are what you would expect from an industry that buys almost nothing locally except concrete, security and landscaping. Its forward linkages are, by construction, non-local: the entire premise of Eastern Data, Western Computing is that the value-added services consuming the compute stay two thousand kilometres east. Its consumption linkage is capped by a payroll of dozens. And its fiscal linkage is the one variable that policy can actually set, which is precisely why competition between jurisdictions tends to bid it towards zero.

This is not an argument that data centres are bad. It is an argument that they are a particular category of thing, and that the category has a well-documented behaviour which the promotional literature systematically ignores. Guizhou did not misunderstand data centres. It understood them as a growth engine, which they are, and hoped they would also be a development engine, which they largely are not.

The Racks That Nobody Rented

The second Guizhou problem is that a good deal of the capital did not even deliver the compute.

In March 2025, MIT Technology Review reported that of the more than five hundred data centre projects announced across China in 2023 and 2024, at least 150 had been completed by the end of 2024, and that local publications were reporting up to 80 per cent of new computing capacity sitting idle. GPU rental prices collapsed accordingly: an eight-GPU Nvidia H100 server that had commanded around 180,000 yuan a month fell to about 75,000.

The DSET researchers Angela Glowacki and Cartus Bo-Xiang You, writing for the Australian Strategic Policy Institute's Strategist in May 2026, mapped the same problem onto the western build specifically. By 2024, 633 hyperscale and large data centres had been built and made operational under Eastern Data, Western Computing, lifting national computing capacity to 268 exaflops. Some western facilities, they wrote, sit empty, with utilisation rates as low as 20 to 30 per cent, a far cry from the original policy goal of more than 60 per cent. Beijing has since restated that all data centres should run at no less than 60 per cent utilisation, and that no new large or super-large facilities should be built in cities where existing ones operate below 50 per cent.

The reasons are mundane and instructive. Remote regions lacked the fibre-optic cables needed to move large volumes of data in real time, forcing operators to spend more on transmission than the model assumed. Renewable curtailment in the western region still exceeds 30 per cent, which undercuts the cheap-clean-power premise. Many facilities were built on the assumption that state-owned enterprises and government agencies would buy the compute, and that demand did not fully arrive. Inter-governmental competition produced speculative overbuilding, and more than a hundred state-backed projects have been scrapped in the past eighteen months against eleven cancellations in the whole of 2023.

Guian, to be fair, is at the better end of this distribution. Local reporting in August 2026 put first-half electricity consumption growth in the new area at 33.7 per cent year on year, on a total of 3.12 billion kilowatt-hours through late June, with big data operations alone consuming 1.93 billion of that, up 52.2 per cent. Guian is busy. But an idle rack and a busy rack impose the same debt service, and a province cannot know which it has bought until several years after it has paid.

Who Actually Pays the Bill

If the benefits are concentrated at the parcel and diffuse to nothing by five kilometres, the costs run in the opposite direction. They start at the parcel and travel.

A 2026 arXiv preprint by Danbo Chen, Zijun Zhou, Yongyang Cai, Jiahong Qin, Ani Katchova and Lei Chen models this directly, coupling language-model analysis of corporate compute plans with energy-system modelling. It projects that electricity consumption by the six largest AI firms will rise from roughly 118 terawatt hours in 2024 to between 239 and 295 terawatt hours by 2030, about one per cent of global power demand, with more than 90 per cent of new capacity landing in North America, Western Europe and Asia-Pacific. Crucially, the burden is not evenly distributed. The authors construct a Power Stress Index and find values above 0.25 in Oregon, Virginia and Ireland, while diversified grids in Texas and Japan absorb the load more comfortably. Their conclusion is that AI infrastructure has become a structural component of power-system dynamics rather than a marginal load, which means it now has to be planned for rather than merely connected.

Water follows the same logic. A preprint by Yuelin Han, Pengfei Li, Adam Wierman and Shaolei Ren, revised in March 2026, estimates that if 2024 water-use intensity persists, US data centres could collectively require between 697 and 1,451 million gallons per day by 2030, comparable to New York City's entire supply. Even assuming aggressive efficiency gains of 10 per cent a year, the range is 227 to 604 million gallons daily. Associated public water infrastructure costs reach roughly ten billion dollars, rising to fifty-eight billion under high-growth scenarios. Their central observation is the one that matters here: these impacts are highly concentrated on communities hosting data centres. The compute is national. The reservoir is not.

Memphis has become the American shorthand for what concentration looks like when it goes wrong. The NAACP, the Southern Environmental Law Center and Earthjustice sued xAI in April 2026 over the operation of 27 unpermitted methane gas turbines in Southaven, Mississippi, effectively a power plant assembled to feed the Colossus 2 facility. This is an airshed where Shelby County in Tennessee and DeSoto County in Mississippi have both received an F grade for ozone from the American Lung Association. The plaintiffs include residents of the Whitehaven and Boxtown neighbourhoods of South Memphis, downwind. The Department of Justice has since intervened on xAI's side. The turbine count, meanwhile, never stopped rising. The plaintiffs went back to court on 6 May 2026 seeking an emergency order to halt operations, and the installations continued regardless: by mid-July, correspondence between xAI's environmental consultant and Mississippi regulators, obtained by Reuters through a public records request, documented 59 unpermitted turbines, at least 57 of them at Southaven, roughly double the number the company had publicly acknowledged. The resolution took the most direct form available: Mississippi's Permit Board had already approved a permanent 1.2 gigawatt plant of 41 turbines on the same ground in March 2026, a month before the suit was filed, and on 31 July xAI agreed a schedule to strip the temporary machines out of the Stanton Road site, beginning in August 2026 and finishing by July 2027. The fight over unpermitted temporary turbines has been settled by making the power station permanent and lawful in the same airshed, breathed by the same people, which is the tell that it was never really about permits. Whatever the litigation concludes, the geography of the dispute is the point: the model is trained everywhere and the generation sits in one postcode, with permission now to stay there.

What Happens When You Ask

There is a final piece of the research picture, and it is a strange one, which is what makes it interesting.

In a preprint first posted in November 2025, Zhifeng Wu, Yuelin Han and Shaolei Ren asked whether large language models could stand in for community consultation on data centre projects. They built a framework that polls AI agents, prompted with local demographic and geographic context, on how they would respond to a proposed facility, and compared the output against real human survey data. The agents identified water usage and utility bills as the dominant concerns and tax revenue as the principal perceived benefit. Responses varied meaningfully depending on which model was used and where the hypothetical project sat. And, notably, the synthetic responses aligned substantially with findings from actual human surveys. The authors propose it as an efficient early-stage instrument for folding neighbourhood perspectives into siting decisions before the plans harden.

You can read that finding two ways, and both are uncomfortable. The optimistic reading is that we now have a cheap way to anticipate what a community will object to, months before the first hearing, at a stage when the design can still change. The bleak reading is that the industry has arrived at a technique for simulating consent, and that the reason such a technique is attractive is that the genuine article is expensive, slow and increasingly likely to say no.

The Guizhou villagers quoted by AFP were not polled, synthetically or otherwise. They were asked a question by a passing reporter and answered it honestly, which is a different exercise from being consulted before a decision. Nothing in the Eastern Data, Western Computing framework required anyone in Guian to be asked whether the mountain should become a campus, and it is worth being clear that this is not solely a feature of the Chinese system. Across the United States, the same decision is routinely taken under non-disclosure agreements and by-right zoning, and communities learn what has been approved after the approval.

Note also what the agents converged on. Water. Bills. Tax revenue. Not wages. Not careers. Not the long-run transformation of the local labour market. Even a synthetic public, prompted to reason about a data centre, does not appear to expect it to be a jobs programme. The expectation gap that Guizhou is living through is largely one that promoters created and that residents, given a moment to think, do not fully share.

What Shu Peihua Is Describing

So return to the shopkeeper on the road outside Huawei's clock tower, because nothing in the preceding sections makes what Shu Peihua said untrue.

The road is real. In a karst province where a mountain village might once have been two hours from a trunk route, a dual carriageway built to carry transformers and chilled water plant is a permanent improvement to the lives of everyone along it. The universities are real; a local vendor told AFP that two had been established since the facilities went in. The customers are real, and so is Li Xixiu's point that people find work nearby now instead of boarding a train to Guangdong. Guizhou has been one of the great exporters of migrant labour in modern China, with all the social cost that implies, and a family that stays together because a parent can get a security job or a canteen job or a fit-out job forty minutes from home has received something that does not show up in a wage-growth ranking.

What the evidence says is narrower and harder. It says that these gains are the consumption linkage of a construction boom plus the ordinary agglomeration of a new-town development, and that they are front-loaded. It says the operational phase which follows will not employ many people, will not raise local salaries measurably, will not spawn a cluster of firms that were not already coming, and will not, on the American evidence, move rents, migration or retail spending much either. It says the electricity, water and land are consumed locally while the value of what they produce is realised somewhere with better weather and higher salaries. And in Guizhou's case, it says the province financed the entry ticket with debt that now ranks second-worst in the country relative to revenue, against assets that may be running at 20 to 30 per cent utilisation.

None of that argues for refusing the data centre. It argues for pricing it honestly, and for noticing which linkage is doing the work. The fiscal one is the only channel a host can reliably control, and it is the one that inter-jurisdictional competition destroys first. There are policy shapes that hold onto it: recurring taxation of the equipment rather than one-off land revenue, as in Loudoun; subsidies capped per permanent job, as Good Jobs First proposes; published utilisation and load data so that a province can tell a productive asset from a monument; large-load tariffs that make the operator, not the household, pay for the substation; and sunset clauses that return the abatement when the promised employment does not materialise.

The alternative is what the resource curse literature has been documenting for sixty years in copper and oil, now rendered in reinforced concrete and immersion cooling. Something enormous arrives. The output figures move. The mountain gets a road, and then the road gets quiet, and the ledger that recorded the growth turns out never to have been the ledger that measured the prosperity.

Shu Peihua is right that it used to be a barren mountain. The question Guizhou has yet to answer, and that a hundred counties from Virginia to Kildare are asking in their own accents, is what a mountain becomes when the thing built on it needs the mountain far more than the mountain needs it.

References and Sources

  1. China goes rural with data centres in quest to power AI – AFP, via The Star, 19 August 2026
  2. Abundant electricity isn't enough: China's overbuilt AI computing power is underused – Angela Glowacki and Cartus Bo-Xiang You, The Strategist, Australian Strategic Policy Institute, 6 May 2026
  3. The Local Economic Impact of Data Centers – Dany Bahar and Greg C. Wright, working paper, 7 August 2026
  4. Concentrated siting of AI data centers drives regional power-system stress under rising global compute demand – Danbo Chen, Zijun Zhou, Yongyang Cai, Jiahong Qin, Ani Katchova and Lei Chen, arXiv preprint 2604.06198
  5. Small Bottle, Big Pipe: Quantifying and Addressing the Impact of Data Centers on Public Water Systems – Yuelin Han, Pengfei Li, Adam Wierman and Shaolei Ren, arXiv preprint 2603.02705
  6. What AI Speaks for Your Community: Polling AI Agents for Public Opinion on Data Center Projects – Zhifeng Wu, Yuelin Han and Shaolei Ren, arXiv preprint 2511.22037
  7. China built hundreds of AI data centers to catch the AI boom. Now many stand unused. – MIT Technology Review, 26 March 2025
  8. Western regions to gain computing power – China Daily, 17 March 2022
  9. Big data key to high-tech development – China Daily, 30 May 2023
  10. China's debt-ridden Guizhou faces reckoning after years of splashing out on pricey projects – South China Morning Post
  11. AI and Data Centers Surge in Guian as First-Half Electricity Demand Jumps 33 Percent – ChinaTechNews, 20 August 2026
  12. Chinese government plans data center capacity reseller network amid overbuild concerns – Data Center Dynamics
  13. Data Centres Metered Electricity Consumption 2025, Key Findings – Central Statistics Office, Ireland
  14. Government Statement on the Role of Data Centres in Ireland's Enterprise Strategy – Department of Enterprise, Trade and Employment, Ireland, June 2018
  15. CRU Publishes its Decision on New Electricity Connection Policy for Data Centres – Commission for Regulation of Utilities, Ireland
  16. Loudoun County, Virginia: The Heart of the Data-Center Boom – Judge Glock, City Journal, 26 April 2026
  17. The Impact of Data Centers on Virginia's State and Local Economies, 6th Biennial Report – Mangum Economics for the Northern Virginia Technology Council, February 2026
  18. Money Lost to the Cloud: How Data Centers Benefit from State and Local Government Subsidies – Good Jobs First
  19. A Generalized Linkage Approach to Development, with Special Reference to Staples – Albert O. Hirschman, 1977, reprinted in The Essential Hirschman, Princeton University Press
  20. Extractive Industries: Optimizing Value Retention in Host Countries – United Nations Conference on Trade and Development
  21. Illegal Pollution from Data Center Power Plants Shouldn't Harm Our Communities. We're Suing xAI. – Earthjustice
  22. Pollution from Musk's unpermitted xAI power project hits hardest in Black communities – Disha Raychaudhuri and Valerie Volcovici, Reuters, via The Star, 14 July 2026
  23. State sets dates to retire temporary xAI turbines but allows some to go past original deadline – Mississippi Today, 31 July 2026
  24. China's north cleans up its power mix as the south lags – Centre for Research on Energy and Clean Air, 19 March 2025
  25. How Many Jobs Does a Data Center Create? – Data Center Geeks, compiling the Uptime Institute 2024 Global Data Center Survey

Tim Green

Tim Green UK-based Systems Theorist & Independent Technology Writer

Tim explores the intersections of artificial intelligence, decentralised cognition, and posthuman ethics. His work, published at smarterarticles.co.uk, challenges dominant narratives of technological progress while proposing interdisciplinary frameworks for collective intelligence and digital stewardship.

His writing has been featured on Ground News and shared by independent researchers across both academic and technological communities.

ORCID: 0009-0002-0156-9795 Email: tim@smarterarticles.co.uk

Listen to the free weekly SmarterArticles Podcast

Discuss...

A pug flying an aeroplane. That was the image, generated with an off-the-shelf model by a photographer who posts under the handle Horshack, and in September 2025 he persuaded a Nikon Z6III to certify it as a photograph.

The method was not sophisticated. He encoded the synthetic image into Nikon's raw NEF container, then used the camera's multiple exposure mode to graft it onto the skeleton of a legitimate capture. The camera, running firmware version 2.00 released on 27 August 2025, duly wrapped the result in a cryptographic manifest conforming to the specification published by the Coalition for Content Provenance and Authenticity. Anyone inspecting the file would have found a valid signature attesting that a real Nikon camera had recorded the scene. Nikon confirmed the fault on 4 September, suspended its Authenticity Service, and later notified users that every certificate it had issued would be revoked. Early adopters who had spent the fortnight signing their work found their proofs retrospectively voided.

The more instructive part came next. Horshack pointed out that Nikon could not fully fix this alone, because the default behaviour of the widely used C2PA validation tools is not to check whether a signing certificate has been revoked. The revocation-checking code exists in the software development kit; it is simply not switched on by default. He filed a GitHub issue asking that the reference command-line tool change that. Until it does, a revoked certificate and a valid one look identical to most software that inspects them.

Hold that sequence in mind, because it contains the whole argument in miniature. Manufacturing a convincing fake cost one person an afternoon and no money. Certifying that something is real required a camera manufacturer, a firmware pipeline, a certificate authority, an industry consortium, a specification, a conformance programme and a public trust list — and it failed anyway, in a way only the consortium could repair.

This is the asymmetry that will define the next decade of the internet, and almost every optimistic account of the “trust economy” walks straight past it.

The Statistic Everyone Cites Cannot Prove Its Own Provenance

Writing in Modern Diplomacy on 16 August 2026, Yang Xite, a research fellow at ANBOUND, set out the case with unusual clarity. “Information used to be scarce, and access to it was a source of power,” he wrote. “Today, the problem is almost the opposite; there is too much information floating online.” His conclusion: “As the cost of producing information approaches zero, the value of finding, checking, filtering, and trusting information rises. The scarce resource is no longer content itself. It is confidence in the content.” And therefore: “In a market where consumers and businesses are increasingly exposed to fraud, proof of authenticity can itself become a product.”

That is correct as far as it goes. The trouble starts with the number recruited to support it.

You will have seen the claim that ninety per cent of online content will be AI-generated by 2026. It appears in policy briefings, keynotes and vendor decks, usually attributed to Europol, whose 2022 report Facing reality? Law enforcement and the challenge of deepfakes states that “experts estimate that as much as 90% of online content may be synthetically generated by 2026”.

In January 2023, Jesse Walker at Reason traced the chain. Europol's “experts” resolve to a single citation: Nina Schick's 2020 book Deepfakes: The Coming Infocalypse. Schick's source was one named individual — Victor Riparbelli, chief executive of Synthesia, a company that sells synthetic video. What Riparbelli said was that synthetic video may account for up to ninety per cent of all video content in as little as three to five years. A forecast about one medium, made by a vendor with a direct commercial interest in that medium's growth, generalised into a claim about the entire internet, laundered through a book into a police agency report and from there into the newspaper of record.

The most widely cited datum about information pollution is itself a specimen of information pollution. It is not a measurement. It never was.

The measurements that do exist tell a duller, more useful story. In October 2025 the search analytics firm Graphite published an analysis by Gregory Druck, Jose Luis Paredes, Bevin Benson and Ethan Smith covering forty-three thousand English-language URLs drawn from CommonCrawl between January 2020 and May 2025. They reported their error rates: a 4.2 per cent false positive rate against pre-ChatGPT articles, and a 0.6 per cent false negative rate against known GPT-4o output. AI-written articles overtook human-written ones in November 2024 and stood at just over half the sample by May 2025. Crucially, the proportion has been broadly stable for about a year. The wave crested; it did not keep rising toward ninety.

Graphite also noted something that cuts against the panic: these articles largely do not surface in Google or ChatGPT results. NewsGuard, tracking the sharper end of the phenomenon, had identified 3,749 AI content farm sites across sixteen languages as of 23 June 2026, more than double its count a year earlier. A real and growing problem. Not ninety per cent of the internet.

This matters beyond pedantry. Build a verification regime on the premise that nearly everything is fake and you will build something that treats unverified as guilty by default. Which is exactly what is being built.

What the Economics Paper Actually Says Is Worse

The arXiv paper usually invoked alongside the ninety per cent figure does not contain it. “The Economics of Information Pollution in the Age of AI: General Equilibrium, Welfare, and Policy Design”, submitted by Yukun Zhang and Tianyang Zhang on 17 September 2025 and revised on 4 January 2026, is a theoretical model rather than an empirical survey. It contains no census of what fraction of the web is synthetic.

What it does contain is a sharper claim. The authors argue that large language models represent “a fundamental shock to the economics of information production” because they collapse the marginal cost of generation asymmetrically — driving the cost of low-quality synthetic content towards zero while leaving high-quality production costly. Formally, AI substitutes for labour in low-quality output but complements it in high-quality creation. From this they derive a unique “Polluted Information Equilibrium”, and attribute its inefficiency to a threefold market failure: a production externality, a platform governance failure, and an information commons externality.

That framework explains something a raw percentage cannot. The problem is not the volume of synthetic content. It is that the cost gap between producing a plausible lie and producing a verified truth has widened, and every actor's incentives now point downhill. Zhang and Zhang propose an Information Pollution Index and an adaptive governance regime keyed to real-time readings — a proposal presupposing an instrument nobody operates.

“AI, Digital Platforms, and the New Systemic Risk” by Philipp Hacker, Lilian Edwards and Atoosa Kasirzadeh, submitted on 22 September 2025 and revised on 23 May 2026, comes at the same territory from the governance side. Drawing on finance, complex systems theory, climate science and cybersecurity, the authors argue that recent European legislation invokes systemic risk while relying on narrow or ambiguous characterisations of it. The Digital Services Act, they find, identifies systemic risk better than the more recent AI Act, and harms such as discrimination at scale and systematic hallucinations “may not fall under current legal definitions, given the AI Act's focus on frontier model capabilities”.

Put the two together and the diagnosis is precise. Generation has become cheap in a way structurally biased towards the low-quality end, and the legal instruments built to respond target frontier model capabilities rather than population-level degradation of the information commons. Neither paper says ninety per cent of the internet is fake. Both say something more troubling: the machinery for noticing and pricing the damage does not exist.

Seeing Stopped Being Enough Somewhere Around Last March

The abstraction becomes concrete fastest in medicine, where the readers are the most highly trained image analysts in the world.

On 24 March 2026, Radiology published a study led by Mickael Tordjman, a post-doctoral fellow at the Icahn School of Medicine at Mount Sinai, testing whether radiologists could distinguish synthetic X-rays from real ones. Seventeen radiologists at twelve centres across six countries assessed 264 radiographs, half authentic and half generated, using both a general-purpose commercial model and RoentGen, a diffusion model built at Stanford for chest radiography.

When the radiologists did not know the study's true purpose, only forty-one per cent spontaneously flagged the synthetic images. Told explicitly that the dataset was half fake, their mean accuracy reached seventy-five per cent, with individual performance ranging from fifty-eight to ninety-two per cent. Experience made no difference: the researchers found no correlation between years in practice and detection accuracy.

Read that again. Specialists, forewarned, missed roughly a quarter of fabricated medical images. The clinical implications are obvious — fraudulent insurance claims, manufactured litigation evidence, falsified trial data — but the epistemic implication is larger. A chest X-ray was, until recently, close to a paradigm case of self-authenticating evidence. It is not any more, and no amount of professional training closes the gap.

The financial sector has the loss figures. On 7 April 2026 the FBI's Internet Crime Complaint Center published its 2025 report: 20.9 billion dollars in reported losses, up twenty-six per cent from 16.6 billion, across 1,008,597 complaints — the first time annual complaints have exceeded one million in the centre's twenty-five year history. People aged sixty and over filed 201,266 complaints and accounted for roughly 7.7 billion dollars of the losses. For the first time, the report carried a section on AI-facilitated fraud: more than twenty-two thousand complaints and nearly 893 million dollars.

The archetypal case remains the engineering firm Arup, whose Hong Kong finance employee authorised fifteen transfers totalling roughly 25.6 million dollars in a single day after a video conference in which every executive on the call was a deepfake assembled from publicly available footage. Hong Kong police disclosed the incident in February 2024; Arup confirmed itself as the victim in May. Deloitte's Center for Financial Services, in an analysis published on 29 May 2024 by Satish Lalchand, Val Srinivas, Brendan Maggiore and Joshua Henderson, projected that generative AI could push United States fraud losses to forty billion dollars by 2027, from 12.3 billion in 2023. That last figure is a projection and should be read as one. The IC3 numbers are not.

Detection Is the Losing Side of an Arms Race It Cannot Leave

The instinctive response is to build a detector. The evidence says this will not work, and the best evidence is a benchmark designed specifically to test the claim.

Deepfake-Eval-2024, assembled by Nuria Alina Chandra, Oren Etzioni and eleven colleagues and first posted in March 2025, collected deepfakes actually circulating in the wild during 2024 — forty-five hours of video, 56.5 hours of audio and 1,975 images, drawn from eighty-eight different websites in fifty-two languages. The team then ran state-of-the-art open-source detectors against it.

Performance collapsed. Against the academic benchmarks on which these models report their headline accuracy, the area under the curve fell by fifty per cent for video, forty-eight per cent for audio and forty-five per cent for images. Commercial detectors and fine-tuned models did better, but still did not match human forensic analysts. The authors' conclusion is blunt: academic benchmarks are out of date and unrepresentative of real-world deepfakes.

This is not a temporary engineering shortfall. Detection is structurally the trailing party. A detector is trained on generators that already exist, and any published detector becomes a differentiable objective for the next generation to optimise against. Compression, re-encoding, screenshotting and platform transcoding degrade exactly the statistical residue detectors look for. Every step of ordinary internet distribution is, incidentally, an evasion technique.

Which means the question cannot be answered by looking harder at the artefact. You have to ask where it came from. And the moment you do, you have changed the subject from what is true to who is vouching — and created a market.

Provenance Changes the Question From What Is True to Who Signed

The Coalition for Content Provenance and Authenticity is the serious attempt at that change. Its output, Content Credentials, is a cryptographically signed manifest travelling with a file, recording what device or software produced it and what edits were applied. The specification reached version 2.4 in April 2026 and is progressing through ISO as standard 22144, still an unpublished draft.

The steering committee is the tell. It comprises Adobe, Amazon, the BBC, Google, Meta, Microsoft, OpenAI, Publicis Groupe, Sony, TikTok and Truepic — the world's largest advertising conglomerate, the dominant mobile and desktop platform vendors, the two largest generative model providers, three major distribution platforms, one public service broadcaster and one verification vendor. A reasonable roster for building an interoperable standard. Also a list of the organisations that will decide what counts as authentic.

Before assessing who should hold that power, be precise about what the technology proves. It proves that a manifest was signed by a key traceable to a listed certificate, and that the file has not changed since. It does not prove the manifest's assertions are true. Photograph an AI-generated image displayed on a screen using a C2PA-enabled Leica and you get a cryptographically impeccable credential attesting that a real camera captured a real scene, which it did — the scene was a monitor. Nor does provenance say anything about context: an authentic photograph from one conflict, correctly signed, remains authentic when captioned as coming from another.

The independent security assessment is worse than the conceptual critique. On 27 April 2026, Enis Golaszewski, Neal Krawetz, Alan T. Sherman, Edward Zieglar and seven colleagues published what they describe as the first comprehensive independent security analysis of C2PA, including the first formal-methods analysis of its core protocols. Their finding: “the current C2PA specifications fail to achieve their claimed security goals”. Their recommendation is stark. C2PA “should not yet be relied upon for high-stakes uses such as financial disclosures, journalism, or legal evidence” — a fairly comprehensive list of the uses for which it is being promoted.

The Certificate Authority Is the Chokepoint

Here the asymmetry stops being conceptual and becomes an access-control list.

To emit a Content Credential validators will trust, you need a signing certificate from a certification authority on the C2PA Trust List. To obtain one, your product must have passed the C2PA Conformance Program, which assesses it against the specification, a certificate policy and a set of security requirements. The interim trust list that held the ecosystem together in its early years was frozen on 1 January 2026; the formal programme replacing it opened in mid-2025 and remains in early enrolment.

Consider what this means in practice. SSL.com, a conformant certification authority, introduced a free tier in June 2026 offering one Level 1 claim signing certificate valid for a year plus ten thousand trusted timestamps. Generous — except that applicants must present “a valid C2PA conformance record ID”. The certificate is free; eligibility to hold one is not something an individual can obtain at all. Conformance is a process applied to products and organisations, with legal agreements, security assessments and identity validation attached. A freelance photographer in Nairobi cannot pass it. A witness with a phone cannot pass it.

The World Privacy Forum, in a technical review published on 3 September 2025 by Kate Kaye and Pam Dixon, put the governance objection with precision: “the criteria used for inclusion on these lists, or who the arbiters deciding which entities are considered known, and by extension, trustworthy have not been made public”. The same review noted that C2PA's own harms modelling documentation concedes that “loss of control over personal information and enforced suppression of speech are possible through use of C2PA”, and that identity specifications were moved out of the core standard into a separate working group in January 2024, with identity claims aggregators now bridging government and commercial identity systems into provenance workflows.

So the architecture, stated plainly: generating a convincing fake requires a consumer graphics card or a free web account. Generating a credential that a verifier will honour requires membership of, or accreditation by, a consortium whose selection criteria are not published. One of those capabilities has been democratised. The other has been enclosed.

Every Platform Reads the Evidence and Then Deletes It

Even for the accredited, the infrastructure does not currently deliver.

In October 2025 the Washington Post ran a test that ought to be better known. Reporters produced an AI-generated video carrying Content Credentials and uploaded it to eight major social platforms. Only YouTube surfaced any warning, and that disclosure sat inside a description attached to the clip rather than on the video itself. No platform preserved the Content Credentials data in a form users could access.

The mechanism is mundane. Platforms re-encode nearly every image and video at upload — for file size, for format normalisation, and to strip EXIF data that routinely contains GPS coordinates and device identifiers. Metadata not deliberately carried through the transcode does not survive it. Some platforms do read the manifest before discarding it: TikTok began reading C2PA data in 2024 and applies its own “AI-generated” labels on that basis, adding an invisible watermark in November 2025 precisely because, in its own account, C2PA metadata can be removed when content is re-uploaded or edited.

That is the settled state of affairs. The provenance record is read by the platform, used for the platform's own labelling decision, and deleted from the file that everyone actually sees. Users receive a conclusion, not evidence. Whether to believe the label is a question about the platform, which is where we started.

The proposed remedy is “durable” credentials — soft bindings such as invisible watermarks and perceptual fingerprints that allow a stripped manifest to be recovered from a cloud registry. C2PA's own FAQ describes the approach. It also relocates verification from the file in your hand to a lookup against a database somebody else operates, and makes every verification a queryable event. Google's SynthID, embedded by default across Gemini, Imagen, Lyria and Veo, had marked more than one hundred billion images and videos by Google's own account in May 2026, along with roughly sixty thousand years of audio. Detection has until now run through Google's detector portal and the Gemini app. At the same announcement Google said SynthID detection and C2PA verification were being built into Google Search and into Chrome, so that a reader can right-click an image, or circle it on a phone, and ask whether it was generated with AI. On the same day OpenAI began attaching Google's watermark, alongside a C2PA manifest, to images produced through ChatGPT and its API; ElevenLabs, Kakao and Nvidia announced adoption too.

Verification offered at the point of reading is a real improvement on everything described above. It also changes nothing about the transcode. The credential still does not survive the upload; what the reader consults instead is not the file but a query, and the largest rival model provider has just adopted the same company's mark. The watermark is Google's, the detector is Google's, and the answer to “is this real” is a request to Google.

Brussels Set a Deadline and It Lands on the Compliant

Two weeks ago, on 2 August 2026, Article 50 of the EU AI Act became enforceable. Providers of systems generating synthetic audio, image, video or text must now mark outputs with “effective, reliable, robust and interoperable machine-readable marks”. Deployers of deepfakes must disclose them with a visible or audible label a person can understand without any detection tool — a hidden machine-readable mark does not satisfy the deployer obligation. Systems already on the market have until 2 December 2026 to comply with the marking requirement, and content generated before 2 August 2026 needs no retrospective labelling. Non-compliance carries fines of up to fifteen million euros or three per cent of worldwide annual turnover, whichever is higher.

This survived a serious attempt to slow the whole regime. The Digital Omnibus on AI, Regulation (EU) 2026/1744, entered into force on 27 July 2026 and deferred the high-risk obligations under Annex III to 2 December 2027 and those for product-embedded systems to 2 August 2028. It did not delay Article 50. Spain has gone further, creating a dedicated supervisory agency, AESIA, and legislating penalties of up to thirty-five million euros or seven per cent of global turnover for failure to label.

Now apply the asymmetry. Article 50 binds providers placing systems on the EU market and deployers operating within its jurisdiction — companies with legal departments, European entities and revenue worth three per cent of. It binds OpenAI, Google, Adobe and Meta. It does not bind an open-weight model running on a laptop in a jurisdiction with no equivalent statute, which is where a growing share of genuinely harmful material is produced. The fraudsters who took 893 million dollars from Americans last year were not weighing their exposure to European administrative fines.

The predictable effect is that compliant, well-resourced, largely benign generation becomes marked, while non-compliant, adversarial generation stays unmarked. Over time that inverts the signal. If the honest are labelled and the dishonest are not, then “unmarked” ceases to mean “probably human” and starts to mean “outside the regime” — a category containing both the criminal and the ordinary person with a phone and no accreditation.

Hacker, Edwards and Kasirzadeh anticipated the structural version of this. Their framework identifies four levels of AI-related systemic risk and warns that the AI Act's orientation towards frontier model capabilities leaves population-level harms outside its definitions. Marking obligations regulate the emitter. The pollution they were meant to address is a property of the commons.

Ten Dollars a Title Is What Being Believed Now Costs

The optimistic version of the trust economy holds that a premium will emerge on genuinely human work. It is worth looking at what such a premium looks like once someone actually builds one.

The Authors Guild operates a certification called Human Authored. A book qualifies if its text was written by one or more humans, with an exception for spelling and grammar tools; using AI for research, brainstorming or indexing does not disqualify a title. The programme opened to Guild members in 2025 and, as of March 2026, to any author published in the United States. Certification costs ten dollars per title for non-members and requires identity verification through a third-party service. Both the fee and the identity check are waived for Guild members.

Two features deserve attention. First, the Guild does not vet manuscripts for AI content before certifying them, and says so, on the reasonable grounds that no reliable detection method exists. Enforcement rests on self-certification, a licensing agreement and community reporting. What is certified is not that a human wrote the book. It is that an identified person has made a contractual assertion and is exposed to consequences if it is false. That is accountability, which is genuinely valuable — and a different product from truth.

Second, look at the price structure. Ten dollars and an identity check for outsiders; free and no identity check for members. That is a small, benign, entirely defensible example of exactly the mechanism at issue. Membership confers the presumption of good faith. Non-membership must purchase it, and surrender identifying data to do so.

Scale that logic to every claim anyone makes online and the phrase “authenticity premium” reveals what it describes. A premium is a price. Someone pays it and someone collects it. Yang Xite's suggestion that “people with established reputations, professional track records, and a history of making reliable judgments may become more valuable precisely because they are accountable for what they say” is almost certainly right, and not obviously good news. It describes a market in which credibility accrues to those who already have standing, and everyone else must buy accreditation from an intermediary or be discounted. A premium on being believed is, in distributional terms, a tax on being doubted. It falls hardest on those least able to pay it and with the weakest prior reputation to trade on — which is to say, most people.

Proof of Personhood Turns Out to Be a Subscription Business

The same shape appears in identity infrastructure, at greater scale and with fewer safeguards.

World, the network built by Tools for Humanity and co-founded by Sam Altman, verifies humanness by scanning irises with a device called the Orb, issuing a World ID that lets a person prove uniqueness without revealing identity. By April 2026, nearly eighteen million people had verified at an Orb. During 2026 the company has pivoted towards enterprise, expanding verification points into retail and introducing fees for applications that check a user's humanity, while keeping verification free for the end user. Reported partners include Tinder and Reddit.

Read that business model carefully. Users supply biometrics free of charge; relying parties pay for the right to trust them. The asset being monetised is the population of verified humans, and the revenue accrues to whoever operates the verification layer. This is not a criticism of the cryptography, which is thoughtful. It is an observation about who ends up holding the register of real people, and on what terms.

The public-sector alternative is further along than most people realise and less far along than its deadline requires. Under the revised eIDAS regulation, all twenty-seven EU member states must make a European Digital Identity Wallet available to citizens by 24 December 2026. Denmark's AltID went into production on 3 June 2026, the first member-state wallet to go live, and passed 230,000 downloads within seven weeks, though it is not yet certified; France Identité and Italy's IT-Wallet are live national services on the same track, and Germany has said its state-issued wallet will arrive on 2 January 2027, after the deadline. Across the twenty-seven the rollout is uneven: fewer than half the member states are on track, the Netherlands among others has signalled it will not meet the deadline, and Bulgaria has not begun work on a state wallet at all. A wallet issued by a state is a materially better arrangement than one issued by a venture-funded company, and it is arriving late into a vacuum commercial providers are filling now.

Meanwhile the enforcement apparatus for synthetic commercial speech is thin. The Federal Trade Commission's rule on consumer reviews and testimonials took effect on 21 October 2024, expressly covering AI-generated reviews, with civil penalties per violation running to tens of thousands of dollars. In December 2025 the agency sent warning letters, later posted publicly, and in January 2026 established a dedicated AI enforcement unit. Set that against NewsGuard's 3,749 content farms and the scale problem answers itself. Enforcement is retail; generation is wholesale.

Absence of Proof Is Becoming Proof of Absence

Everything above converges on a single failure mode, and the World Privacy Forum named it before it arrived.

Provenance systems classify. Content carrying a valid signature is marked as such; content without one is rendered “unknown” or “invalid”. In an interface, on a phone, glanced at for a second and a half, those categories collapse into one: not verified, therefore suspect.

Consider who is systematically unverified. A protester filming police conduct on a five-year-old handset with no provenance capability. A whistleblower who must strip metadata to survive. A journalist in an authoritarian state for whom a cryptographic link between a file and a device is a targeting package. A photographer in a country with no conformant certificate authority. The World Privacy Forum found precisely this risk: that the trust model penalises creators lacking C2PA adoption and disadvantages marginalised and at-risk communities — the very groups the technology's designers set out to protect.

There is a second-order effect. When authenticity becomes a live question, real evidence loses force. The liar's dividend — dismissing genuine recordings as fabricated — is available to anyone whose accuser lacks a certificate, and that accuser is usually the one with less money. Provenance infrastructure does not merely fail to help such people. It hands their opponents a rhetorical instrument: where are the credentials?

Nor is the burden evenly distributed geographically. The conformance programme, the certificate authorities, the steering committee and the standards process are concentrated in the United States, western Europe and Japan. The trust list is the operative boundary of the verifiable internet, and it was drawn by eleven organisations, none accountable to an electorate, according to criteria that have not been published.

What Would Actually Redistribute the Burden

None of this argues for abandoning provenance. Content Credentials are useful, and a signed chain of custody from a wire agency's camera to a newspaper's front page beats nothing. The argument is that provenance as currently constituted transfers the cost of doubt onto the people least able to bear it, and that this is a design choice rather than a law of nature.

Four changes would alter the distribution.

Make absence neutral by design. Interfaces should be prohibited from rendering unsigned content as suspect. The presence of a credential is information; its absence is not. This is a user-experience rule with constitutional weight, and the cheapest intervention available.

Open the trust list. If the conformance criteria and the identity of the arbiters remain unpublished, a private consortium is exercising a public function without accountability. Publishing the criteria, the appeals process and the removal grounds costs nothing. An individual pathway to signing — accreditation of persons, not only products and organisations, with the cost borne publicly — is the harder and more necessary step.

Fix revocation before scaling. The Nikon episode was contained because Horshack disclosed it, and could not be fully remediated because validators do not check revocation by default. A provenance ecosystem in which a revoked certificate validates is worse than no ecosystem, because it manufactures unearned confidence. Golaszewski and colleagues have supplied the formal analysis; the specification should be treated as pre-deployment until their findings are addressed.

Regulate the reader, not only the writer. Article 50 obliges emitters to mark. It says nothing about the platforms that strip those marks in transit. A distribution-side duty to preserve and display provenance data, enforceable against very large platforms under the Digital Services Act, would make the marking obligation meaningful. At present the EU compels the creation of evidence and permits its destruction seconds later.

The pug flying the aeroplane was never really about Nikon. It was about the discovery that the certificate and the thing certified had come apart, and that only the certifier could put them back together. That is the position ordinary people are being moved into across every domain where truth must now be demonstrated rather than observed. The ability to fabricate has been handed to everybody. The ability to be believed is being metered, priced and administered by a small number of hardware manufacturers, platforms, model providers and identity vendors, on terms they have not published, with an appeals process that does not exist.

Confidence is indeed becoming the scarce resource. Whether that is a market opportunity or a civil liberties emergency turns not on whether verification gets built, but on whether the people who cannot afford it are treated as unproven or as liars.

References

  1. Yang Xite, “Information Pollution and the Rise of the Trust Economy,” Modern Diplomacy, 16 August 2026. https://moderndiplomacy.eu/2026/08/16/information-pollution-and-the-rise-of-the-trust-economy/
  2. Yukun Zhang and Tianyang Zhang, “The Economics of Information Pollution in the Age of AI: General Equilibrium, Welfare, and Policy Design,” arXiv:2509.13729, 17 September 2025, revised 4 January 2026. https://arxiv.org/abs/2509.13729
  3. Philipp Hacker, Lilian Edwards and Atoosa Kasirzadeh, “AI, Digital Platforms, and the New Systemic Risk,” arXiv:2509.17878, 22 September 2025, revised 23 May 2026. https://arxiv.org/abs/2509.17878
  4. Jeremy Gray, “Nikon Can't Fully Solve the Z6 III's C2PA Problems Alone,” PetaPixel, 22 September 2025. https://petapixel.com/2025/09/22/nikon-cant-fully-solve-the-z6-iiis-c2pa-problems-alone/
  5. Europol Innovation Lab, “Facing reality? Law enforcement and the challenge of deepfakes,” Europol, 2022. https://www.europol.europa.eu/cms/sites/default/files/documents/Europol_Innovation_Lab_Facing_Reality_Law_Enforcement_And_The_Challenge_Of_Deepfakes.pdf
  6. Jesse Walker, “That time we tried to make sense of a statistic in a 'New York Times' story on deepfakes,” Reason, 27 January 2023. https://reason.com/2023/01/27/that-time-we-tried-to-make-sense-of-a-statistic-in-a-new-york-times-story-on-deepfakes/
  7. Gregory Druck, Jose Luis Paredes, Bevin Benson and Ethan Smith, “More Articles Are Now Created by AI Than Humans,” Graphite, October 2025. https://graphite.io/five-percent/more-articles-are-now-created-by-ai-than-humans
  8. NewsGuard, “Tracking AI-enabled Misinformation: 3,749 AI Content Farm sites (and Counting),” NewsGuard AI Tracking Center, 23 June 2026. https://www.newsguardtech.com/special-reports/ai-tracking-center/
  9. Mickael Tordjman et al., “The Rise of Deepfake Medical Imaging: Radiologists' Diagnostic Accuracy in Detecting ChatGPT-generated Radiographs,” Radiology, 24 March 2026. https://pubs.rsna.org/doi/10.1148/radiol.252094
  10. Federal Bureau of Investigation, “2025 Internet Crime Report,” Internet Crime Complaint Center, 7 April 2026. https://www.fbi.gov/file-repository/2025_ic3report.pdf/view
  11. Heather Chen and Kathleen Magramo, “Arup revealed as victim of $25 million deepfake scam involving Hong Kong employee,” CNN Business, 16 May 2024. https://www.cnn.com/2024/05/16/tech/arup-deepfake-scam-loss-hong-kong-intl-hnk
  12. Satish Lalchand, Val Srinivas, Brendan Maggiore and Joshua Henderson, “Generative AI is expected to magnify the risk of deepfakes and other fraud in banking,” Deloitte Center for Financial Services, 29 May 2024. https://www.deloitte.com/us/en/insights/industry/financial-services/deepfake-banking-fraud-risk-on-the-rise.html
  13. Nuria Alina Chandra, Hannah Lee, Ryan Murtfeldt, Lin Qiu, Arnab Karmakar, Emmanuel Tanumihardja, Kevin Farhat, Ben Caffee, Changyeon Lee, Jongwook Choi, Sejin Paik, Aerin Kim and Oren Etzioni, “Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024,” arXiv:2503.02857, 4 March 2025, revised 27 May 2026. https://arxiv.org/abs/2503.02857
  14. Coalition for Content Provenance and Authenticity, “C2PA,” c2pa.org, accessed 16 August 2026. https://c2pa.org/
  15. Enis Golaszewski, Neal Krawetz, Alan T. Sherman, Edward Zieglar, Sai K. Matukumalli, Roberto Yus, Carson L. Kegley, Michael Barthel, William Bowman, Bharg Barot and Kaur Kullman, “Verifying Provenance of Digital Media: Why the C2PA Specifications Fall Short,” arXiv:2604.24890, 27 April 2026. https://arxiv.org/abs/2604.24890
  16. Kate Kaye and Pam Dixon, “Privacy, Identity and Trust in C2PA: A Technical Review and Analysis of the C2PA Digital Media Provenance Framework,” World Privacy Forum, 3 September 2025. https://worldprivacyforum.org/posts/privacy-identity-and-trust-in-c2pa/
  17. SSL.com, “C2PA Certificates for Media Authenticity,” ssl.com, accessed 16 August 2026. https://www.ssl.com/products/content-authenticity/content-credentials/c2pa/
  18. Washington Post, “Tests show top social platforms don't disclose markers on AI videos,” 22 October 2025. https://www.washingtonpost.com/technology/2025/10/22/ai-deepfake-sora-platforms-c2pa/
  19. Google DeepMind, “SynthID,” deepmind.google, accessed 16 August 2026. https://deepmind.google/models/synthid/
  20. European Commission, “Transparency obligations under Article 50 of the AI Act,” Shaping Europe's Digital Future, 2026. https://digital-strategy.ec.europa.eu/en/faqs/transparency-obligations-under-article-50-ai-act
  21. Gibson Dunn, “EU AI Act Omnibus Agreement — Postponed High-Risk Deadlines and Other Key Changes,” gibsondunn.com, 2026. https://www.gibsondunn.com/eu-ai-act-omnibus-agreement-postponed-high-risk-deadlines-and-other-key-changes/
  22. The Authors Guild, “Human Authored Certification,” authorsguild.org, accessed 16 August 2026. https://authorsguild.org/human-authored/
  23. Biometric Update, “World shifts from crypto identity experiment to enterprise proof-of-humanity,” biometricupdate.com, June 2026. https://www.biometricupdate.com/202606/world-shifts-from-crypto-identity-experiment-to-enterprise-proof-of-humanity
  24. European Commission, “European Digital Identity Wallet,” ec.europa.eu, accessed 16 August 2026. https://ec.europa.eu/digital-building-blocks/sites/display/EUDIGITALIDENTITYWALLET/EU+Digital+Identity+Wallet+Home
  25. DLA Piper, “FTC's 2026 enforcement approach to fake reviews takes shape: Takeaways for companies,” dlapiper.com, July 2026. https://www.dlapiper.com/en-us/insights/publications/2026/07/ftcs-2026-enforcement-approach-to-fake-reviews-takes-shape-takeaways-for-companies

Tim Green

Tim Green UK-based Systems Theorist & Independent Technology Writer

Tim explores the intersections of artificial intelligence, decentralised cognition, and posthuman ethics. His work, published at smarterarticles.co.uk, challenges dominant narratives of technological progress while proposing interdisciplinary frameworks for collective intelligence and digital stewardship.

His writing has been featured on Ground News and shared by independent researchers across both academic and technological communities.

ORCID: 0009-0002-0156-9795 Email: tim@smarterarticles.co.uk

Listen to the free weekly SmarterArticles Podcast

Discuss...

Enter your email to subscribe to updates.