Watermarking just alters the pseudorandom number generator. If "I like turtles" was previously the response to your prompt with probability 100%, it will still be so. This is why watermarking is only effective for long strings of text
It's like the sudden change of a language style and its verbosity didn't happen recently.
To random words you pick and provide a sufficient amount of text to vary with random number without losing its meaning you need a text with high entropy.
Nothing about watermarking would require padding the response length with pseudo-intelligible Claudese. Regular filler would work fine.
Also, it would probably provide higher entropy to write normal human-sounding English instead of reusing a repetitive grab bag of load-bearing phrases. This theory doesn't really make any sense.
> As far as I know, anthropic aren't intrinsically motivated by watermarking (if anything it hurts sales, and seems indifferent to safety(?)) they're simply doing it to fulfill the EU obligations.
They are. They want to reduce the amount of LLM generated text they feed into their next model training.
Also, how would you watermark a sentence with just 3 words for an example? This exactly why it became so verbose.
That would be a terrible tradeoff. The ship has already sailed and a lot of public AI content will not be their own. Deliberately making their product worse to reduce identifiability of AI inputs by 25% just doesn't sound worth it to me. Is that what you would pick if you were in charge of anthropic and wanted to maximise the company's product?
And what wisdom do you think they would be missing if unable to distinguish three word written pieces? Keep in mind that most sources are not inherently trustworthy just because they rate as human written, too. You need some other way to rate text in all cases.
I know Adel from a professional community, but I've followed his work for a while and it's the real thing. I would defenitely recommend him if you are looking into programmer who is deep into compilers and RUST. On top of that he is also agile, curious guy, who had a chance to play with various technologies.
I think the best compromise would be to get the best of two words. By default perform bound checks, but have a compiler flag which skips it. Might broke many programs written with default behaviour in mind, but allow perform additional optimizations.
this is exactly what julia does. boundschecks are default on, and there are compiler flags --- either locally, via the `@inbounds` macro, or globally with `--check-bounds=no`--- to disable them
Hi ziotom!
I wonder about you work in 3D Cifford Algebras. May you share some links to the research you do? I also have interest in this topic I research on my own.
Just in case if you don't want to disclose your name my email is northzen@gmail.com
To machine. It just easier to be polite by default than split our language into two forms "I speak to human" and "I speak to machine". Because the chat interface is really close to what we see when we speak to human. Well, exactly the same.
> It's about the fact that AI can be used to extract value from other artists' work without consent, and then out-compete them on volume by flooding the marketplace.
I didn't even think about the analogy to sampling (and the prior controversy) but that is an even better analogy. Ultimately, the different between what's creative re-use and what's a ripoff is a matter of how skillfully it's done and there's a lot of controversy in the middle!
This style provides a high entropy basis distribution, so they can from a bigger pool to pick from and phrases to watermark the sentence.
You have much bigger variation of this idiotic phrases and words, which states a simple fact in that sophisticated and twisted manner.
reply