Rendered at 08:21:53 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
timpera 19 hours ago [-]
I recently tried Pangram, created an account, wrote a few lines about how my day went, and it was flagged as likely AI-assisted. It clearly doesn't work.
It's especially bad that they keep insisting that it works very well, because thousands of people will probably end up falsely accused of AI usage as a result.
skippyfish 18 hours ago [-]
Can you share your experiment, so that we can draw our own conclusions? I've been paying for Pangram for a long time and have not seen it misbehave even once. I've seen many people attack it, but when I looked more closely, their claims often evidently boiled down to "I want plausible deniability for my own use of LLMs" - i.e., they were clearly posting AI-generated content on social media and just didn't like the possibility of being called out.
I'm not saying that's you, but since a proof is easy to produce, it would be nice if you could share.
amrit3128 18 hours ago [-]
I think most native English speakers are doing well. The one who will face issues are non native ones. I knew English as my third language since about 10 years but I've only truly become somewhat proficient in it in the past 5 or so years, thus it was heavily influenced by LLMs. So who's at fault for this? Me, for speaking like an LLM simply due to my circumstances? Am I now supposed to learn "human" english?
aleph_minus_one 18 hours ago [-]
Relevant concerning your point:
> I'm Kenyan. I Don't Write Like ChatGPT. ChatGPT Writes Like Me.
From what I've seen and read it regularly flags native English writing as AI too. Sometimes even changing a few words is enough for it to change the verdict.
It's not command of the english language it flags (for native vs non native to matter), it's stylistic mannerisms, some of which happen to also exist in some direct translation of some foreign patterns for some languages.
But it detects them in a crude manner, missing obvious AI slop, and misflagging human texts.
ano-ther 18 hours ago [-]
I‘ve seen reports of markdown-tags (##) flipping a pangram result from human to machine (with human written text).
But I don’t have a license to confirm.
coldtea 16 hours ago [-]
>I've been paying for Pangram for a long time and have not seen it misbehave even once.
How would you know if you just take whatever slop verdict is serves as correct?
phoghed 18 hours ago [-]
Read the OP you goofball, there’s examples right there in it.
skippyfish 18 hours ago [-]
I'm asking the parent about their claim. I already have an opinion about the parent post, in which the author is claiming to have not used an LLM to author a 2025 essay that's illustrated with completely unnecessary AI art. Although one doesn't imply the other, let's say that this diminishes my willingness to believe that claim at face value.
embedding-shape 18 hours ago [-]
What about the piece the author wrote in 2017 and because of an addition of 23% AI generated text, Pangram labelled it as 100% AI written?
Ariarule 17 hours ago [-]
More generally, Freddie deBoer has been writing in publicly available works and published for so long that he's obviously more credible for how any previous work was put together than what appears _itself_ to be a hallucination-prone AI system. Granted only one person knows for sure, but the idea that Pangram may have actually been correct all along doesn't pass a laugh test for me.
Do people not recognize names nor search for them before making insinuations? It's not like OP is a brand-new pseudonymous essay on a default-template blog with one or two other posts in the history: https://en.wikipedia.org/wiki/Fredrik_deBoer
skippyfish 14 hours ago [-]
That's not relevant at all. I'm asking for human-written text that gets flagged by Pangram as AI-generated, not an artificially-constructed case that contains AI text (and was explained by the founder of the website as essentially a block resolution issue).
> Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes.
> When disagreeing, please reply to the argument instead of calling names. "That is idiotic; 1 + 1 is 2, not 3" can be shortened to "1 + 1 is 2, not 3."
> Please don't comment on whether someone read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that".
Wowfunhappy 16 hours ago [-]
The problem is likely that you only wrote a few sentences. Pangram clearly needs more text than that to work, which makes sense.
Pangram should refuse to work with small amounts of text. (Well, I know it already does, but the threshold should clearly be significantly higher.)
ThrowawayTestr 14 hours ago [-]
If Pangram worked the AI companies would be using it to improve their models
johnfn 17 hours ago [-]
This issue is, like, exactly what the response from Pangram is talking about.
btown 14 hours ago [-]
It frustrates me that AI detection for student essays effectively works as a protection racket: if a student wants to write an essay without AI assistance and reliably get credit for doing so, they still need to pay a company like Pangram for an individual account to ensure their own work isn't accidentally flagged (especially if checking multiple drafts a day).
And even this isn't perfect, nor is it guaranteed that enterprise and individual accounts are tuned the same way. So students also need to proactively use audit/keystroke logging systems to protect themselves against accusations, which creates a type of "panopticon" on one's early/ephemeral drafts, including language of frustration (who among us hasn't typed curses into an unsaved draft at some point?), that can massively stifle creative thought. And if an institution provides such a tool, their centralized access simply worsens the "panopticon" characteristics.
There's no easy solution, here, sadly.
runako 16 hours ago [-]
This product does not even have a plausible theory of how it could work.
LLM-generated text does not carry a watermark or other identifying marks. The "theory" is that an LLM trained on human writing, to mimic human writing, can be distinguished from actual human writing in under 100 words.
Notably the first diagram on the research overview page (https://www.pangram.com/research/how-it-works) shows feedback for "misclassified human examples." This is a category error; Pangram will not find out when it has misclassified text in the wild, except in rare cases. Only the "licensed human-written text" in its training data can be used as feedback.
Scams like Pangram also cause real harms, mostly because laypeople do not understand that what is being offered is not possible. Pangram advertises 99.98% accuracy, and they pitch it as a tool for teachers and universities. Translated: if a college like University of Alabama rolled this out, you could expect ~40 students to have their lives upended by this snake oil, every year. (And how can one even prove that an allegation is false, that they did write a given text?) And this is the best case, using the number on Pangram's homepage.
Smaug123 2 hours ago [-]
Claude, a general-purpose model, can identify me, personally with stylometry in about 200 words. Is it really such a stretch to believe it’s possible for a special-purpose model to identify the ten or so main LLMs crossed with the fifty or so main styles people gave them write in?
Catloafdev 18 hours ago [-]
I'm not sure I'd use the term 'brittle' to describe snake oil.
achileas 20 hours ago [-]
Can something be broken that never actually worked?
pixl97 17 hours ago [-]
A real "not even wrong" moment.
sscaryterry 20 hours ago [-]
That is indeed the question. The little bit I tried Pangram, it gave me unpredictable and surprising results.
13 hours ago [-]
embedding-shape 20 hours ago [-]
> If I want to induce a false positive in TSA's airport scanner, I can put a gun-shaped object in my bag. If I want to induce a false positive in Waymo's stop sign detection system, I can paint my own sign and put it up on a pole.
Holy strawman-batman, not only does the founder of Pangram not have a proper response to the actual criticisms, he feels the need to completely make up very different situations to try to illustrate some completely different point... I guess good job of the founder to engage at all, as deBoer does bring up a lot of valid points and criticisms of why people really shouldn't rely on "tools" like Pangram, too bad the founder failed completely at addressing the more serious points, and instead just chose to say "Well, there will be false-positives, what can you do?".
cgio 19 hours ago [-]
Indeed, that’s not about false positives predominantly. At least for me the key concern is if it’s giving a percentage when it looks like it calculates a Boolean and therefore implies more nuanced analysis than it does. Which explains the nesting behaviour too.
ameliaquining 18 hours ago [-]
I'm confused, what are the "actual criticisms" you think weren't responded to?
phoghed 18 hours ago [-]
The Russian nesting doll of 100% human and 100% AI is a big one.
johnfn 17 hours ago [-]
Pangram said a bunch of reasonable things and responded very coherently to the point and you skipped over it and misinterpreted this one section.
coldtea 16 hours ago [-]
What he said amounts to "you're holding it wrong".
johnfn 12 hours ago [-]
The response says "I also believe we can do better". How is that in any way "you're holding it wrong"?
coldtea 17 hours ago [-]
Those detectors are even more snake oil shit than some early "AI companies" which had people in like India do the actual work.
And the obvious absolute test that would prove its shit - millions of books and posts written pre-AI would unfortunately be in its training already with some date associated, and thus it would not really detect them, it would just know "x text, written pre 2020".
18 hours ago [-]
650 18 hours ago [-]
Adding another anecdote, but Pangram flagged an essay I wrote last week as AI-assisted when it wasn't. I take Pangram results with a grain of salt now.
phoghed 18 hours ago [-]
Even if everyone tacitly agrees Pangram is bullshit, there’s perhaps an opportunity to evolve it into a Yelp-like racket anyway.
It's especially bad that they keep insisting that it works very well, because thousands of people will probably end up falsely accused of AI usage as a result.
I'm not saying that's you, but since a proof is easy to produce, it would be nice if you could share.
> I'm Kenyan. I Don't Write Like ChatGPT. ChatGPT Writes Like Me.
> https://marcusolang.substack.com/p/im-kenyan-i-dont-write-li...
HN discussion:
> https://news.ycombinator.com/item?id=46273466
It's not command of the english language it flags (for native vs non native to matter), it's stylistic mannerisms, some of which happen to also exist in some direct translation of some foreign patterns for some languages.
But it detects them in a crude manner, missing obvious AI slop, and misflagging human texts.
But I don’t have a license to confirm.
How would you know if you just take whatever slop verdict is serves as correct?
Do people not recognize names nor search for them before making insinuations? It's not like OP is a brand-new pseudonymous essay on a default-template blog with one or two other posts in the history: https://en.wikipedia.org/wiki/Fredrik_deBoer
> Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes.
> When disagreeing, please reply to the argument instead of calling names. "That is idiotic; 1 + 1 is 2, not 3" can be shortened to "1 + 1 is 2, not 3."
> Please don't comment on whether someone read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that".
Pangram should refuse to work with small amounts of text. (Well, I know it already does, but the threshold should clearly be significantly higher.)
And even this isn't perfect, nor is it guaranteed that enterprise and individual accounts are tuned the same way. So students also need to proactively use audit/keystroke logging systems to protect themselves against accusations, which creates a type of "panopticon" on one's early/ephemeral drafts, including language of frustration (who among us hasn't typed curses into an unsaved draft at some point?), that can massively stifle creative thought. And if an institution provides such a tool, their centralized access simply worsens the "panopticon" characteristics.
There's no easy solution, here, sadly.
LLM-generated text does not carry a watermark or other identifying marks. The "theory" is that an LLM trained on human writing, to mimic human writing, can be distinguished from actual human writing in under 100 words.
Notably the first diagram on the research overview page (https://www.pangram.com/research/how-it-works) shows feedback for "misclassified human examples." This is a category error; Pangram will not find out when it has misclassified text in the wild, except in rare cases. Only the "licensed human-written text" in its training data can be used as feedback.
Scams like Pangram also cause real harms, mostly because laypeople do not understand that what is being offered is not possible. Pangram advertises 99.98% accuracy, and they pitch it as a tool for teachers and universities. Translated: if a college like University of Alabama rolled this out, you could expect ~40 students to have their lives upended by this snake oil, every year. (And how can one even prove that an allegation is false, that they did write a given text?) And this is the best case, using the number on Pangram's homepage.
Holy strawman-batman, not only does the founder of Pangram not have a proper response to the actual criticisms, he feels the need to completely make up very different situations to try to illustrate some completely different point... I guess good job of the founder to engage at all, as deBoer does bring up a lot of valid points and criticisms of why people really shouldn't rely on "tools" like Pangram, too bad the founder failed completely at addressing the more serious points, and instead just chose to say "Well, there will be false-positives, what can you do?".
And the obvious absolute test that would prove its shit - millions of books and posts written pre-AI would unfortunately be in its training already with some date associated, and thus it would not really detect them, it would just know "x text, written pre 2020".