In July 2025, Deloitte Australia delivered a 237-page report to the country’s Department of Employment and Workplace Relations, priced at roughly A$440,000. It reviewed the automated system Australia uses to penalize welfare recipients. The report cited a Federal Court judgment. The quote was invented. It cited academic papers on the welfare system. Several of them did not exist.
A Sydney University researcher, Chris Rudge, found the fabrications and told the press the document was full of fabricated references. Deloitte confirmed the errors, quietly published a corrected version disclosing that Azure OpenAI had been used to help write it, and refunded part of its fee. The firm maintained its conclusions still held. Nobody at Deloitte, apparently, had checked whether the sources in a report bearing the firm’s name actually existed.
That is the story worth sitting with before any NGO decides how much editorial work to hand to a language model. Not because the technology is useless. It plainly is not. But because the failure mode isn’t obvious sloppiness — it’s fluent, well-formatted, entirely plausible text that happens to be wrong, and the people paid to catch that are, increasingly, not catching it.
The citation problem is getting worse, not better
Deloitte’s report wasn’t an isolated embarrassment. A Columbia University team led by Maxim Topaz audited more than 2.5 million papers on PubMed Central after Topaz caught a fabricated citation an AI tool had slipped into his own manuscript. What they found, published in the Lancet, was a steep and accelerating trend:
| Period | Rate of papers with at least one fabricated citation |
|---|---|
| 2023 | 1 in 2,828 |
| 2025 | 1 in 458 |
| First seven weeks of 2026 | 1 in 277 |
That’s roughly a twelvefold increase in three years, and the researchers noted the fabrications weren’t sloppy or obviously wrong. They used correct formatting, real researchers’ names, and plausible publication dates, which is exactly why they got through peer review.
News writing shows a similar pattern. A recent analysis comparing 100 AI-labeled articles against 100 human-written ones found AI-labeled pieces were 8.2 times more likely to contain a hallucinated claim — a fabricated quote, a wrong statistic, an event dated to the wrong month. None of that is a hypothetical risk for an NGO. It’s precisely the kind of detail — a funding figure, a policy citation, a quote from a partner organization — that a communications team can’t afford to get wrong, because getting it wrong once is usually enough to lose a donor’s or a journalist’s trust for good.
Sourcing is where AI writing quietly breaks down
The mechanism is worth understanding rather than just fearing. A language model doesn’t retrieve a citation the way a person pulling a book off a shelf does. It predicts what a citation-shaped string of text should look like, based on patterns in its training data. Most of the time that produces something real, because most citation patterns in its training data were real. Sometimes it produces something that looks exactly as real and simply isn’t — a paper with a plausible title, a plausible author, a plausible journal, and no actual existence.
This is precisely the kind of error a careful human editor is built to catch, because catching it means doing something a language model structurally doesn’t do on its own: stopping to ask whether a claim is actually true rather than whether it merely sounds right. A source can be checked by opening the link. A quote can be checked against a transcript. A statistic can be checked against the original dataset. None of that is glamorous work, and all of it is exactly what separates an NGO communications piece a funder can rely on from one that quietly erodes the organization’s credibility the first time someone checks a footnote.
Design has the same problem in a different costume
Text hallucinations get the headlines, but the same pattern shows up in visual work, and NGOs have particular reasons to be careful here. A study of 171 AI-generated images published by voluntary organizations, from major international development charities down to small grassroots groups, found two things worth flagging. First, disclosure was inconsistent — more than one in ten images carried no AI credit at all. Second, and more troubling, the imagery tended to reproduce the same reductive, dehumanizing tropes about poverty that the sector has spent decades trying to move away from, rather than offering anything genuinely new. When researchers looked at the public comments under six charity campaigns using AI imagery, only a small fraction actually discussed the cause the charity was raising money for. Most were arguments about whether the images were fake.
That backlash isn’t limited to the nonprofit sector. When Vogue ran a Guess advertisement using AI-generated models in its August 2025 issue, it disclosed the AI use in a footnote and still triggered a wave of public criticism, with readers pointing to the magazine’s own history with photographers like Irving Penn and Lee Miller as the standard AI imagery failed to meet. Disclosure alone didn’t settle the question of trust. Something about knowing an image wasn’t real changed how the audience read the rest of the message around it.
For an NGO, whose entire relationship with a donor rests on being believed, that’s not a minor branding footnote. It’s close to the whole business model.
What a human editor is actually doing
None of this is an argument that AI tools have no place in NGO communications. Drafting, summarizing, restructuring, translating a rough outline into something readable — all of that is faster with a language model in the loop, and no serious communications team is going to give that up. The argument is narrower: the further a piece of content gets from a human’s final read-through, the more likely it is to carry an error nobody meant to publish. A fabricated statistic in a grant report. A quote attributed to the wrong spokesperson. An image that echoes exactly the stereotype the campaign was trying to avoid.
A human editor checking that work isn’t performing a ritual. They’re doing the one thing the underlying technology cannot yet reliably do for itself: verifying that what sounds true is actually true, and that what looks right actually reflects the organization publishing it. For a sector whose credibility is its entire asset, that check is not overhead. It’s the job.
Sources
- Deloitte to partially refund Australian government for report with apparent AI-generated errors — AP/Fortune
- Deloitte refunds over $60K for report with AI errors, Australian government says — CFO Dive
- AI Blamed For Rise In Fabricated Citations Found In Recent Research Papers — Forbes
- AI use in American newspapers is widespread, uneven, and rarely disclosed — arXiv preprint
- AI-Generated Images in Charities: The Governance Gap — Insights2Outputs
- Vogue US faces backlash over Guess ad featuring AI-generated model — FashionNetwork





