> On that note, one way we can prevent it is to assert that all our content is byte-for-byte identical with the last known trusted stage of what we have produced
That doesn't help with things like the typical use of SynthID where the spymarking is done by the same process generating the content, so there is never a clean comparator. (It also wouldn't be useful anytime it is inplemented as part of a transformation—compression, etc. —step, for the same reason.)
This has been going on for a while with Facebook. They seem to embed custom metadata tags so that images shared outside the platform can be traced back:
https://stackoverflow.com/questions/31120222/iptc-metadata-a...
Stop using links instead of words. Your comment is literally unreadable without going on to other websites.
:D
Repo here https://github.com/sutt/innocuous. It works with last year's llama.cpp. Check out the "Use Cases" and "How it works" sections in the readme if you're interested.
You just need to address three questions:
- who controls what information is going in? (that is, what is the process by which the tech companies who control all the tech are using it)
- who controls what information is coming out? (that is, is the steganographic format open enough that anyone can read it, or does it depend on having a key)
- what legal regulation is this subject to? (does sneaking individuals name and address into their photographs incur you massive GDPR liabilities when it is discovered?)
Note that there's a widespread precedent: https://en.wikipedia.org/wiki/Printer_tracking_dots
Or... add some noise. Just align the last bit of every pixel channel with a random bit sequence - and Bob's your uncle.
First the low-end laptops and phones (and probably later, most of them) will incorporate some low-level driver that is constantly scanning for these and passing them to a helper app to phone home. I assume this is something Apple will, to their credit, refuse to do[1] but I don't think other OEMs will have any qualms based on what they already do with their TVs.
[1] (though they don't do this kind of thing out of altruism, but because their cash cow is app store rents and fat hardware margins, not third-party advertising.)
I still use IRC on the daily, but someone mentions Discord at least once a month on there, and sometimes tries to whisk people away to that side. There are also people hooking up LLMs to IRC bots and joining them to channels without permission, then when you complain or kick/ban their bot you're somehow treated as the rude one. It's very hard to entirely get away from all the crap anymore.
Like it choose between "winding" and "curving" but there are many uses of curving that probably can't be replaced with "winding" like "her gently curving thighs" with "her gently winding thighs"
But I'm sure there are some intricacies I don't understand. Anyway, very cool website, thanks for sharing it~!
Of course, the absence of a watermark/spymark doesn't prove that the source wasn't AI generated. But the absence provides evidence that it wasn't.
AI companies should simply use watermarks in the responsible sense: they should indicate that the material was AI generated, not include personal information in it.
We also recently had this with LG spy-TVs. Cars here in the EU also spy on people, allegedly to show how alert they are. Perhaps they sneakily upload that information somewhere ... Facebook also has the spy-glasses now. People getting angry about Flock-spy-cameras.
It seems we are now in the age of spying of everyone at all times. Future spying will be done via even smaller devices.
Digital Rights Management isn't just about restricting what you can do with the content, which is often futile anyway. Another way companies can manage their rights is by making sure pirates are properly identified and caught.
It's a shame however, how low quality and vibecoded the live examples are. The first example says "Toy example; not SynthID.", the second one is a generic spectrogram and the third one has an identification space too small to be useful (173 in decimal). I was hoping to see more realistic scenarios to learn how these new watermarks are being applied, instead of generic steganography.
You can’t definitively prove the absence of a watermark. You can only prove the watermark is there. Once you do prove it’s there, the thing that carries the watermark changes in some way — it is “burned” or tainted?
There must be value in having a visible vs an invisible watermark, or in declaring that a work is watermarked without revealing the hidden mark, or having two marks — one that is publicly verifiable and another that is hidden?
If the process itself can be defeated through adding entropy (or more generally by revealing the watermark algorithm) then is that not security through obscurity, which is to say it is a one-shot rather than a general system that is doomed to become obsolete over time?
Something feels off about a technology based on being hidden but whose only value is in being revealed but I feel dumb for not being able to be more specific about what feels wrong! It could simply be that anyone who can verify the presence of the watermark also now has a tool to tell them when they’ve successfully scrubbed the watermark off the work, so the verify tool has to be kept secret which in turn limits its usefulness.
How do you spymark text that someone else wrote? You can't change the words or they'd notice
Text the whistleblower only reports on, well, if they got it from a computer system, there's already precedent of altering word choices, typos and punctuation in e-mails and memos to create unique per-recipient or per-recipient-group versions, which allows companies to trace leaked transcripts reported by press back to source of the leak.
Tech like SynthID I see a net positive especially since it doesn't degrade text quality. I dream about a browser extension running at all times that makes text more translucent based on the confidence of LLM writing[0].
This article's suggestion of using it to unmask whistleblowers is very interesting and not something I'd thought about though. Still not convinced that spymark is a better name though.
[0]: Sean Goedecke's Deckard is close but it would rather invisible than bright red https://www.seangoedecke.com/deckard/
https://www.abc.net.au/news/2026-08-18/what-happens-if-prope...
Might be technically more accurate, but it’s an inscrutable name with no chance of proliferation beyond technical people. If the goal is to rally people to your position, you need a name people can identify.
1. be okay with sending their prompt to an AI company?
2. need to publish something AI generated?
I don't think its fair to say that metadata on apps will be safely removable in the future.
Seems like a wet dream for DRM with lots of possible uses that may be considered bad, but there's also some potential to have it be used to better control what you publish and own, enabling you to exactly steer where and how your content should/can be consume.
On the other hand, no matter how robust that solution is, inevitably someone will come up with a way to bypass it, strip them out, etc - so would it really be useful in the long run?
I recall they had separately watermarked versions of these to make it easier to figure out how things were being leaked.
The "tracking" bit is kind of nefarious, but that can be removed as a concern if the thing that is being tracked is agents, not users.
Not a single example provided of anything that could be honestly called "your work", just a bizarre attempt to stigmatize accurate detection of genAI output.
What was that PG bit about "submarining"?
It's not "your work" it's the bloody AI's work! That's the whole point.
If it is text, copying text alone and not the file will it not remove it? Massage the text with Ai and vola spymark gone, don't you think?
No. The mark is hidden in the word choices.
See the demo in the article.
[deleted]
[dead]
[dead]
[dead]
[dead]
Hidden side channels of any kind, not under the user's control, should be looked upon with suspicion.
> > Because I always wanted to coin something. Please don't forget me.
Watermarks are not "spymarks". They're DRM. I wouldn't worry about advertisers tracking conversions. I would worry about the "analog hole" being closed. Think of no longer being able to even photograph your phone screen, because pixels on it carry digital watermark that's robust enough to survive being photographed - I.e. the kind currently used to tag AI generated images - and then every phone and computer refusing to display resulting photo because the app disallowed capturing its pixels.
No, the words could contain steganography. Use links to be safe!
(Edit: hmm, without the "https://" it seems to depend on the browsers ability to recognise a URL.)
"for stenography (link)"
Or to use another HNism
"Stenography[1]"
Those interested could click it, those not could still read the comment.
What will be the next? I will be unable to see the domain of a link on hover/longtap and have to trust random links like on a search engine?
MUTINY against hn!
* the 7-bit quantization + 1-bit noise option not by much, but still, visible
It’s your choice if you believe them or not, I like Apple and I wouldn’t use that feature.
The pretending that this is the same thing, that Apple is sneaking something past you when they’re showing you that they’re trying to do it right is a bad faith argument.
I think you may have an outdated view here
But fuck it, just brand all our brains with "SLA Industries".(fictional dystopian corporation ruling future)
Who, exactly, is the "bt@brand.io" whose only attributable contribution is FUD regarding SynthID (and what are their motives)?
So I’m not writing off tech as a whole just the adtech companies being a lost cause.
TrojanStego: https://arxiv.org/abs/2505.20118 Improvement: https://arxiv.org/abs/2606.09411
The concern is valid, but microphones are far worse. They're simpler, smaller, extremely sensitive to sound and an order of magnitude cheaper, both the mic itself as well as any spying with it.
It is possible to record voice using few bytes, to send later. It's further possible to transcribe cheaply into text, and analyze said text.
And mics are already everywhere, including in devices that do not need them, as well as speakers that can be rewired by software to act as microphones.
I spent half a day messing around with it and I was very impressed by how robust it is. I couldn't get OpenAI to stop detecting their own SynthID without completely trashing the image.
The closest thing I found to "defeating" SynthID was to put in a normal photograph and ask ChatGPT to make some utterly trivial edit, and then the output got flagged with SynthID even though it is essentially an unmodified photograph.
You absolutely can, but if the process transforms the input, it requires you to understand the transformation, or to use an identical, trusted transformer.
[deleted]
The article calls out watermarks intended to deter counterfeiting as explicitly being not spymarks. Watermarks can tell you about the items marked, not about the person using/creating it.
One is proof of authenticity, other is tracking tool.
I'm not convinced spymark is better than just "invisible watermarks", spymark to my ears sounds designed to be sound very negative when invisible watermarks are not always negative, e.g. the counterfeit bank note example.
The article spent quite a bit of energy explaining why the word choice, seems like you're just ignoring that? Also, bank note watermarks are not invisible. Watermarks are not invisible, as the article (again) took pains to explain. Tech like SynthID I see a net positive especially since it doesn't degrade text quality.
It absolutely does. It constrains high-entropy word choice so it can "store" other things in your text. Your text actually has information removed from it. I dream about a browser extension running at all times that makes text more translucent based on the confidence of LLM writing.
Sounds very much not worth the anti-consumer, anti-privacy aspects which (again) the article explains.Bank notes are mainly protected by things that are hard to create without very specialized machines. You can't rely on anything hidden staying unknown and once you know a steganography scheme you can also control it.
> Tech like SynthID I see a net positive especially since it doesn't degrade text quality. I dream about a browser extension running at all times that makes text more translucent based on the confidence of LLM writing[0].
False confidence is a lot worse than no confidence. If such an extension ever becomes popular, people will take anything not marked as AI as gospel.
I wonder if we could have a real, open and direct democracy. All the models we have right now work via indirect clowns. Then again, looking at how some people vote, perhaps direct democracy can only work if people are clever.
To do politics right take skills, and I don't expect the average mechanic to be better at it than the average politician is at fixing cars. How should I know if we should subsidize organic farming, ban alcohol sale after 8PM, or increase the defense budget? At least in theory, politicians are professionals who deal with these kinds of questions, they are supposed to know the technical and social implications, or find experts to help them if they don't. Some people think they know, and judging by how stupid most of their ideas are, they don't. I don't blame them, it is just not their field, and my ideas are probably just as stupid anyways.
I would. Politicians are pre-selected for people who want to lead and that's the last kind person that should be allowed to.
The problem with these spymarks is that they can be used to include data that's completely invisible to users - even potentially to sophisticated users and the programs that consume the marked files. So while I can make an informed decision whether to share a picture, I may not be informed about any spymarks. Vs a normal watermark that aren't designed to be invisible.
Because they have this pathway only they can access, their key and signature can be trusted. I’m not gonna trust Joe Schmoe’s signature that “No I didn’t use AI” unless I already trust Joe Schmoe (and in which case, he doesn’t need a watermark, I’ll just believe him when he says it).
[dead]
When some processing is required by eu or eu member state law, that processing doesn't need explicit consent.
One could also argue the watermark tracks the generated content, not the person
[dead]
These kids and their foreign contractor run VPN services are naive about depth-charge Steganography. Privacy has been dead for years. =3
Don't worry about it... Have a nice day, =3
Like, a significant portion of the people who were involved in any screwups made in the 2000s -- especially the more senior people who should get most of the blame -- are now retired.
Make decisions based on today and on reality. Apple holding a grudge against Nvidia today, if true (as we have no proof that it factors in, of course), is downright childish.
Apple had the opportunity to take Apple Silicon to the datacenter and become the #1 professional ARM microarchitecture in the world. It was Apple's stupid grudge that gave Nvidia everything they needed to sell thousands of their lazily-made Grace CPU.
Even if you trust them, maybe as an user you can be ok with that. As a non-user who will talk with people wearing Apple Watches, I disagree being recorded and my conversations with the watch owner summarized.
Where do I disagree for that ?
e.g. Johnny's accused of something white collar, did he ever make any prompts that suggest how early on he was aware of {X} and further indicate how he moved to frame it?
That's a requirement that varies by country.
They’re asserting that all the audio is done on device, and the results of encrypted so they can’t access them even from the backups.
Unlike… EVERY… other tech company, it is in Apple’s interest to be privacy-focused.
Even if you just have to believe them, which you do pretty much, they’re the biggest name pushing for privacy in the world right now. They make more on selling devices than they make on ads and behaviors. It’s in their interest to not lie.
Of course, at that point it depends on how much you trust their PCC.
But make no mistake, Apple itself is an ad company and that sets all the incentives that matter.
Oh, really?