Although this looks like a clever approach, a kind of stochastic key, I do not see how this guarantees to distinguish text written by big babble machines versus humans. Humans also have a certain pattern of writing, a given distribution of how some words are more likely to appear than others. How can one tell them really apart?
As an indicator, yeah, might be usable. But I wouldn’t read too much into it before seeing results of a study that runs actual tests.
It’s not about the variation of the words, it’s about the variation of the words from the model baseline.
Like if your word choice was almost the exact same as Claude’s normally, maybe you just talked to them a lot and picked up their phrases like it’s not nothing.
But if you managed to be almost exactly like Claude and yet varied the possible words exactly according to a hidden entropy key, they’d know it was actually Claude with the SymthID-Text watermarking applied, as no human would end up falling into that statistical bucket.
Although this looks like a clever approach, a kind of stochastic key, I do not see how this guarantees to distinguish text written by big babble machines versus humans. Humans also have a certain pattern of writing, a given distribution of how some words are more likely to appear than others. How can one tell them really apart?
As an indicator, yeah, might be usable. But I wouldn’t read too much into it before seeing results of a study that runs actual tests.
It’s not about the variation of the words, it’s about the variation of the words from the model baseline.
Like if your word choice was almost the exact same as Claude’s normally, maybe you just talked to them a lot and picked up their phrases like it’s not nothing.
But if you managed to be almost exactly like Claude and yet varied the possible words exactly according to a hidden entropy key, they’d know it was actually Claude with the SymthID-Text watermarking applied, as no human would end up falling into that statistical bucket.
I thought the article explained that pretty reasonably on a scale of probability and weight. The longer the text, the more reliable the scoring.