When I write True’s stories, I use copy/paste from the source articles for quotes to make absolutely sure I don’t introduce errors with my own typos: I never want to accidentally put the wrong words into anyone’s mouth. This week, that practice revealed something interesting.
Let me start with the story I was writing, since it is ironically germane to what I saw.
What, Again?
Government officials in South Africa drafted a “national A.I. policy” to set the country up as an innovative leader in the field, and to address the ethical, social, and economic challenges related to A.I. Once the draft was released for public comment, several in the public commented — pointing out that many of its references didn’t exist. “The most plausible explanation is that A.I.-generated citations were included without proper verification,” said Communications Minister Solly Malatsi. “This should not have happened.” He ordered the draft to be withdrawn. “This failure is not a mere technical issue but has compromised the integrity and credibility of the draft policy,” he said. The references included articles from real journals such as AI & Society and the Journal of Ethics and Social Philosophy, which confirm the articles cited do not exist. (RC/London Independent) …Click here if you would like me to draft your letter of resignation.
While writing, I could see that there was something different about what I was pasting from what I knew I was copying. I had a suspicion, and because I know ChatGPT “sees” what’s actually pasted, I copied a sample into its prompt box to ask what the London Independent was doing, and why. Its reply was interesting — my suspicion was correct, according to Chatty-G:
The Independent is inserting invisible Unicode control characters — such as zero-width spaces and word joiners — into the article text. Readers never see them on screen, but they survive copy-and-paste operations, subtly altering the text. The technique is increasingly used by publishers as a form of anti-scraping or anti-AI “watermarking”: it can interfere with automated text harvesting, identify copied material, and sometimes help trace the source of reproduced content.
Which makes this particular story especially amusing, since the article in question is about using AI to help draft policies governing AI use — without disclosing it. The modern media ecosystem increasingly resembles an arms race in which everyone involved is using AI while simultaneously trying to keep everyone else’s AI from using their work.
Arms race. Yeah, that’s about right. The news media fiddles while civilization burns.
(Link to my source story at the Independent.)
What I Was Seeing

To be clear, the above screencap is how my editor interpreted what was pasted in. To really see what is there requires something that can turn what is “seen” by machines into something that’s human-readable. For which I used the old Unix utility od, or “octal dump” for the second paragraph shown above:

The irony of their putting scraper poison into quotes that they copied from another platform is not lost on me. I did copy over some other test passages to check them; some had the alterations, which show in a different (and bold-faced) font, some didn’t. I did check the Malatsi post on Twixxer the same way, and did not find the alterations; I thus conclude they were added by the Independent as they are not in the original.
The first time I saw this pattern I thought it was just a glitch, and simply had cGPT clean it up for me. That was a cinch for it to do …which (I’ll avoid saying “ironically” again) shows just how effective the tactic is. I don’t remember if that previous time was at the Independent or elsewhere, but as cG says, “The technique is increasingly used by publishers” — plural — and was probably somewhere else as I don’t source from the Independent very often, even if I also spotted this week’s “headline of the week” while I was there.
When I saw it again while writing the above story, I realized it had to be purposeful, thought I knew why, and got the confirmation from Chatty-G, which I then re-confirmed with other tools to double-check against hallucinations.
Ethical Lapse?
I’m of mixed mind as to whether this amounts to an ethical lapse — that these publications are doing the exact thing I work to avoid: they’re changing quotes, at least in this specific case. It is minor, but it is stretching toward, if not crossing, a line.
It’s extremely common for news operations to write their own versions of a story, hopefully crediting the other news sources where they find it (as I do). I see it all the time while researching stories, and ironically (cough) wrote about that exact thing last week in my Author’s Notes, though I didn’t post that brief note in my blog. Here’s what I wrote:
Classic: A reader suggested one of the stories this week. The URL they found was on MSN. MSN said their source was a radio station, and, happily, provided a link to that. The radio station said their source was a TV station, and again provided a link. You guessed it: the story wasn’t originally reported by the TV station, either: they linked to a story at the local newspaper, which is the one I used for the True treatment.
As much as possible, I prefer to follow the chain to the original on-the-scene journalist who reported the story. Anyone who has played telephone (often called “Chinese whispers” in British countries) knows why. But at least I give kudos to the outlets who not just admitted where they got the info, but linked to it. Very often, outlets don’t.
When the Independent plays Telephone, they purposefully add static to the line, in case someone is listening in. Ethical lapse, or valid business tactic? You tell me: the Comments are open.

P.S.: Why did I slug the story “What, Again?”? It’s a nod to earlier stories about A.I. policy gaffes, such as one about Dawn, a Pakistani newspaper, which included, at the end of a story in a printed edition, “If you want, I can also create an even snappier ‘front-page style’ version with punchy one-line stats and a bold, infographic-ready layout perfect for maximum reader impact. Do you want me to do that next?” Readers recognized that construction. One posted online, “Dawn, Pakistan’s leading newspaper, was caught using ChatGPT despite its strict AI policy.” (“Consider the Source”, 16 November 2025)
– – –
Bad link? Broken image? Other problem on this page? Please Let Me Know using the Help button in the lower right, and thanks.
This page is an example of my style of “Thought-Provoking Entertainment”. This is True is an email newsletter that uses “weird news” as a vehicle to explore the human condition in an entertaining way. If that sounds good, click here to open a subscribe form.
To really support This is True, you’re invited to sign up for a subscription to the much-expanded Premium edition.
Re “Classic”: Another newsletter I read linked to an interesting story. That story offered at the end a link to read “the original story on People”. So I read that, which didn’t have any more info, but linked to their source which was CBS8. The CBS8 story did have some extra details so I was glad I traced it through. But it seemed very lazy reporting on behalf of the first story to credit People as the “original” source when they had a link to their source right there!
Sadly, with the economic pressures on newsrooms, I expect this sort of thing (concomitantly with the proportion of news that is just reposts from other sources) will only increase in the near and medium term.
—
Yes, I see that all the time too. I agree with your conclusion. -rc
To draw a circle, I sometimes see AI “assists” in search results even though I try to turn them off. Many of these now quote sources and, you guessed it, they are rarely the original (or correct) source.
—
Back in Google’s “Don’t Be Evil” days, they insisted they would never become a publisher since that would compete with the sites they were there to index. Next step: get rid of the “Don’t Be Evil” slogan, and here we are. -rc
I suppose you’ve highlighted that the technique is misapplied (to the extent that it works): watermarking your work is legitimate, but not the bits you aren’t pretending are your own work.
Then again, if it’s directed at AI harvesting and not again diligent humans, maybe it’s sensible.
—
Except that, as noted, it doesn’t work. -rc
As an old guy initially versed in WordPerfect, I’m assuming this is somewhat like ‘reveal codes’ mode, which was quite useful and straightforward. On that tool’s demise, I’m stuck in Word where the ‘reveal codes’ tool was far less friendly.
I used to take to keeping the Text Editor handy, and I’d paste into that tool. Because it was very simple, it didn’t know how to handle the hidden codes and ditched them, so re-copying that paste effectively stripped off the hidden junk quite easily.
Notepad kinda accomplishes this task these days. Although sometimes I’ve resorted to the url field for shorter passages!
Unless I got the concept very wrong….
—
You’re pretty much on track. The editor I refer to is Textpad, which has been updated for modern character sets, and thus doesn’t simply strip out “odd” characters. It rather gives an indication of something weird (my first screenshot). The Unix od utility gives a real analysis to show what’s going on behind the scenes, much like WordPerfect’s wonderful “Reveal Codes”, which I still use because there has been no demise of WordPerfect. They don’t update it often (I think the most recent is from 2022). I haven’t felt the need to update it from the 2020 version I use to write TRUE. -rc
Ah, Jim, how much I miss ol’ WordPerfect! Their “reveal codes” feature was absolutely superb. Moreover, when you activated any formatting, it carried from that point until the ending control character code. For example, boldface could carry over several paragraphs or indeed many pages. Same with change of font face, or size, or …. well, you get the point.
MS Word uses the approach of all formatting control characters are repeated anew each paragraph. This greatly increases the data size on the disk (and in file transfer, and in memory usage, and…) and it is all JUST BLOAT.
Sadly, this approach (as far as I know, invented by MS) is now the standard for about every editor. *sigh*
—
At the risk of repeating what I told Jim: there is no reason to “miss ol’ WordPerfect!”, at least if you are on Windows. I use it every day. It’s still maintained. It’s still sold. It’s still great. It’s still orders of magnitude better than Word. -rc
To Randy’s sage advice: When corporate says you use Word over WordPerfect, therein lies the dilemma.
—
Which is just one reason I wanted to “be” The Man. That said, at NASA I was the go-to guy for editing papers because (thanks to Wordperfect) I could get complex math equations set correctly. “No one” that I ever heard of seemed to be able to do it in Word: they would leave space, and then paste them in as graphics. -rc
True or false: “purposefully add static” equals “unethical”.
—
That’s what I asked at the end, yes. And your answer is? -rc
Where I work we routine have developed content for deployment in the Moodle Learning Management System (LMS). Many of the folks involve work with MS Word (*ugh*!) and naturally enough it tosses in lots of control characters as you see in your “anti-AI scraping.” These control characters are harmless in Moodle but they do add to the data size. For example, I routinely find words that are boldface, but the bold on / bold off codes are inserted willy-nilly, sometimes in the middle of a word. I have even seen it were a sentence is boldface but the bold off code precedes every space and the space is then followed by a bold on code. Total waste.
Admittedly, it’s not usually severe but by hand going through and taking out all of the unnecessary control characters has been shown to reduce the overall document size by routinely as much as 5%-15%. (In one extreme case, the bloat was well upwards of 50%). This may not seem to terrible to North Americans on unlimited Internet services, but for students in many locations in Africa and Asia, who purchase every megabyte, this bloat wastes their money. And they don´t even know it.
One trick we have developed is using “paste as plain text”… in many editors, rather than CTRL+V, use CTRL+SHIFT+V. This pastes in the visible characters and strips out all of those annoying control characters.
—
To be sure, this essay doesn’t state that there is never any reason for such characters: they exist for a reason. But I will add that I’ve never seen such in news articles for 31 years, and now, suddenly, they’re there (on some, certainly not all, news sites). Coincidence? Not according to cGPT, as shown. But indeed, Word has terrible code bloat, which is why I won’t use it unless it’s to edit someone else’s files. -rc
What no one has mentioned is the effect of this anti-AI-inspired garbage on legitimate academic works attempting to accurately quote — and properly cite — passages. I often instruct my students – not for AI-related reasons — to use the “Paste as plain text” option, but now it seems there is an additional reason: To be able to paste an accurate quote! I normally encourage the use of cutting and pasting precisely because retyping things often results in quotation errors — especially with poor typists such as myself!
On another note, I have long bemoaned the adoption of Microsoft Word as the academic and business world standard. Even back when I first began using PCs in 1989 I saw WordPerfect as the far superior word processing program, but over the years it became increasingly necessary to convert papers to Word format until I gave up WordPerfect altogether. Sadly, the case in technology is often one of the better-marketed product surviving: VHS vs. Betamax, Adobe vs. Corel, etc.
—
Indeed, but how did you know I still use CorelDraw? 🙂 And this is the last comment I’ll approve about software compatibility: it’s getting us too far from the point of this page. -rc