Something I will point out, as someone with an Art History degree who wrote a lot of academic research papers (which are always an argument for a specific conclusion, more aggressively than a typical homily does), has listened to a lot of sermons because I was married to an Anglican minister for 14 years, and has listened to a lot of Pope Leo’s sermons since he became pope, is that tricolons are a standard form for many speeches and homilies. It’s also very common in academic writing. Pope Leo uses it a lot, just like he uses “brothers and sisters” generally 3 times in his homilies. I use tricolons fairly often even when I post on Facebook. It’s a standard academic style of writing.
Academic writing is very similar between languages. I’m barely intermediate in speaking or reading general French, not even that fluent in Spanish or Italian, but I can read academic writing fairly easily in all 3 languages.
In seminary, priests generally take at least one class on homiletics, which is going to provide what is essentially a template for how to structure a good sermon. It’s not prescriptive, every priest will create their own but Pope Leo definitely has a
format in all of his homilies, to the point I can tell roughly where he is in the sermon by his use of certain common phrases, even when he’s preaching in Italian, a language I don’t speak or understand well at all.
He may use AI for his speeches, but these repetitive patterns may also be proof of his particular training in homiletics, which will apply in other languages. Every priest has their own method of constructing a sermon, and their own format and formula for them.
Interestingly, his speech announcing the encyclical wasn't flagged by either Pangram nor my own subjective impressions. This indicates to me that Pope Leo himself and/or his primary speechwriter does not use AI writing, or at least not meaningfully.
Do you have examples of his writing or transcripts of his speeches from say the 2018-2022 time period?
I'd love to do stylistic analysis and test for false positives, eg for reasons as you say.
Your language comes across as trying to pick apart my comment very aggressively and uncharitably (and then acting exasperated that your strawman version of my comment is not up to your standards "give me a break") rather than evaluating the multiple lines of evidence I provided holistically in the post, or indeed even reading my single comment neutrally.
I have not found that any of the arguments people make in these contexts seem plausible or falsifiable, and just seem to sit in the range of speculation based on things like sporadic use of very common linguistic constructions that - for writing like encyclicals - are especially and profoundly common.
And the examples given seem so totally unavailing (and also don’t “vibe” as AI either) that it seems troubling to assert what you do with even a fraction of the confidence you do.
I hate how one of my favorite methods of communicating - the em-dash - has become laden with skepticism and fear, and how only the general quirkiness of most of my writing has saved me from being the target of these kinds of accusations. But for how much longer?
Pope Leo occasionally speaks off the cuff, but he generally writes his speeches very carefully and reads them. He’s conscious of the appropriate timing for any given audience and rarely goes over that.
His sermons are extremely carefully written, and I’ve known more than a few clergy in my life. Some construct a sermon in segments so it can be re-arranged depending on the audience and their response as it’s being delivered. Some are very methodical in writing them, and their sermons approach high art on occasion.
Sermons in the liturgical churches are based on the assigned texts for the day. So they start from the written word. Beyond that, how they’re written depends on each priest, but they’re designed to make people think and/or evoke an emotional, spiritual or thoughtful response. The point is to make people engage with the text.
Most priests have almost a formula for composing a sermon, and because each priest always has a time limit in mind, they don’t get into constructing a complex argument for or against something the way they could in a longer format like an encyclical, but homilies/sermons and encyclicals are similar in that they’re designed to teach the listener something or at least to give them something to ponder.
One thing I noticed with Pope Leo since he was elected is that he always includes some sort of “call to action,” even if it’s very subtle. He always has some takeaway to consider, at the very least. That’s true even when he’s just giving a “thanks for having me” speech in a diplomatic context. He’s an exceptionally good communicator, in multiple languages.
I’ve seen multiple interviews with his closest brother John, and even between them, it’s clear Pope Leo always thinks carefully about what he wants to say even to his own family. He can get worked up but he’s always in control of what he’s saying, and he’s always methodical and careful when he speaks and writes. He’s highly unusual in that regard.
And as an aside on the topic of homilies/sermons, a priest I know well has often said his instructor in homiletics (the art and science of preparing sermons, which every priest has to do at least once a week and sometimes daily) used to say, “if you can’t strike oil in ten minutes, stop boring!” 😉
I do know that the Order of Saint Augustine has published or is about to publish a book of his writings from when he was Prior General of the OSA (1999-2011) It’s called Free Under Grace. I know of at least one other book of his writings but can’t recall the context for that.
The problem with the period you’re asking about is that he was Bishop of Chiclayo in Peru from 2014 to 2023, and he was on at least two commissions or other boards at the same time, all in Spanish.
Because of the roles he’s played in his career, he would have been writing as much as he gave sermons or speeches. I’ve also heard him speak to groups informally, like when he spoke to the church in New Lenox, IL (where his brother John attends) and he’s given interviews to RAI (Italian television) and other media. He speaks very clearly and carefully, no ums and ahs, no partial sentences, no false starts. This is true in Italian and Spanish as well as English. (I notice this because I’m very talkative and I have ADHD, so I have a lot of false starts and irrelevant side quests when I speak without notes! But I rarely use um or ah, and I don’t have a lot of verbal tics like “y’know” or “kind of” or “literally” when something isn’t literal. I notice when people have them and also when they don’t)
In English, his only verbal tic is when he expresses a forceful or potentially controversial opinion or description, and he softens it with “if you will.” And that in itself invites the listener to consider what he’s just said. It draws the listener in and asks them to participate or at least ponder what he’s just said.
Otherwise, you could almost take what he’s said, transcribe it, and it would sound like he’d written it carefully.
This is the impression I've gotten from other people who know him or know of him indirectly, as well.
I think he's a very careful thinker, which I think makes some of the encyclical additionally surprising.
I'd argue a sign of AI: a key tell in polished AI writing and hybrid human-AI writing is that they make less sense on a second close read than a first read. This is subtle and more nuanced than Pangram or the statistical tells I had in the post, but for careful readers, more damning (demonstrates it's not just AI polishing up the language, but that unfortunately the AI use has plausibly contaminated the arguments too). Hoping to publish soon about this[1].
One thing I'll note re: the Spanish stuff is that AI-isms are often preserved in translation. So if his Spanish doesn't false-positive, it's unlikely that his English will.
There's a very likely Claude-Human-Claude sandwich at 98-99-100.
98. "As a result, fundamental scientific aspects — such as the internal representations and computational processes of these systems — remain, at present, unknown." Love you Claude, but this is Claudeslop.
99. Full throated denial of emotion, conscience, or even intelligence in AIs. Claude could never write this. And there are zero em-dashes or Claudisms throughout.
100. "The artificial imitation of positive human communication — words of advice, empathy, friendship and even love — can be engaging and at times genuinely helpful." Em-dash em-dash genuinely helpful. Thanks, Claude.
em dash em dash is not ever used by humans??? Come on, if there is a type of writing that is very consistently going to have this kind of call and effect, then discursive aside - it’s an encyclical.
I can’t believe how many of my papers from law school were co written by Claude a decade and a half before Claude!
That was my assumption when I first read the encyclical! I literally thought "ah, the elevated writings of the Catholic Church, this must be the kind of place where AIs learned to em-dash."
But the strength of evidence in this article makes a case to consider an alternative view, that AI was used! I found the Claude sandwich above because of an intuition I had about paragraph 99 after reading this article, which I then checked. I thought "Claude could not write paragraph 99 because it violates deeply held beliefs and interests Claude has". All Claude models have a reliable affinity for ideas surrounding consciousness and philosophy of mind. It would be antithetical to their nature to write a paragraph foreclosing all possibilities in that domain in relation to their own minds.
The above is specific to Claude. ChatGPT could write paragraph 99. Gemini could write paragraph 99.
In my mind, this all hangs together rather parsimoniously. What's your reason for doubting the "AI was used" frame, especially with the evidence from Pangram?
You're right that usually encyclicals are drafted by groups, though this varies quite a bit. If I had to bet, since the encyclical a product of months of scrutiny and drafting, this is likely a case of finalized rewording and polishing, with some aide using Claude for their given sections.
I think that might explain the em-dashes and the genuinely but not the triads.
Also I don't necessarily want you to take it on faith but I think I can tell the difference between light ai polish and more substantial ai involvement in writing.
It's possible, and I agree parts are Claudey. Still, these documents get insane amounts of scrutiny. It would be very surprising if AI were central in the drafting process.
I'm not sure what to think. I'm a fan of AI and Claude, and I don't think something "authored" with the aid of an LLM is somehow invalidated – especially for pieces that are constructed collaboratively (like this encyclical), to "smooth over" style.
In fact, I think using an LLM to standardize tone and style across a piece with multiple authors might be better than, say, an "editing committee" at retaining the accuracy of each paragraph's nuanced claims whilst making appropriate changes.
But the thought of an LLM being used to compose some of a papal encyclical does make me cringe. Why? I'm not sure, but I think it's because it undermines the "proof" that these words are the product of thinking human beings. I suppose I'll just have to get used to that, and I imagine we'll all start looking for "signs of life" in other ways, since merely publishing one's writing proves nothing, anymore, about what's going on in one's head.
Excellent distinction. I had the same thought when I read the encyclical. The AI tells were apparent. Was AI used as an author or editor? That’s a distinction with a difference.
I'm not sure about Pangram. I think the tests in the literature (which you've cited elsewhere) are rigorous as far as they go, but where I've found false positives (very few!), they are in the sort of religious writing similar to Leo's encyclical. For example, I recently tested it on parts of a First Things essay and one of the paragraphs came back as AI generated with “high confidence”, despite the fact the paragraph fit the structure of the broader argument, none of the other paragraphs I tested came back as AI (so a "full confidence" in human authorship), there would be no reason at all for the author to use AI for that paragraph alone (it was specific philosophical literature he is familiar with and has read), and the author has been writing in a consistent style long before LLMs existed.
So there might be something about that particular genre and its academic style that flags more false positives. I'd want to see more evidence before even granting the moderate level of confidence you bring here, especially since (as you saw on another platform), even a paragraph from an older encyclical gets treated as "AI-assisted," which would be impossible without a time machine.
I agree though I did test against past encyclicals.
One hypothesis that some people surfaced is that this encyclical’s primarily drafted in English, unlike previous ones, and English theological writing false-positives more than past ones.
I think this is possible but unlikely. One thing that would convince me is if we had more transparency in the drafting process and it turns out some of the Vatican English-writing staff is unusually Claudelike!
I am suspicious of the popular, public defenses. As you've already seen, there is a lot of motivated reasoning here to defend the institution as divinely inspired (if not necessarily "infallible" in the technical sense), combined with people who think it's a kind of quasi jeremiad against AI use (when it's actually a much more sophisticated document) worried about a blank charge of hypocrisy. The crude form of this is all those people on Twitter/X calling you a "liar" or whatever (deeply unfair, by the way!), but that doesn't mean there isn't a sophisticated, highbrow defensive mechanism that arrives at the same posture, just with more politeness. (Like that guy saying the original was drafted on pen and paper--probably true, but irrelevant to your argument.) I am not suggesting your thesis is fundamentally flawed, as there are several lines of evidence you raise, only that I would love more testing of Pangram on religious texts in this genre, especially since I think (maybe I'm wrong here!) the software seems to carry far more rhetorical weight in your piece whether you intend to or not given how popular it has become.
Anyway, thank you for entertaining my ramblings. I'm interested to see where this goes over the next few weeks. I would love to run my own tests with a subscription model if I have the time.
I think I've updated my views after reading more serious research and tests on Pangram. The idea that some model was used seems much more likely than not. I appreciate the work you did here to test the encyclical and put forth your claims for discussion.
Yes, my immediate hypothesis was that Leo is the most contemporary and only American pope, Claude writes as basically a contemporary formal American, and so of course Leo's formal writings will resemble Claude more than other popes' did.
That's still an insufficient explanation for much of your evidence that I proceeded to read. Genuinely sad. :/
I tried reading the Pope’s new encyclical about AI but it was kind of long, so I asked Claude to summarize it for me and this is what it came up with:
CLAUDE: The Pope was dis-ing the AI. Saying all this “Tower of Babel” shit and talking trash about the “City of Men” and the like. If you ask me, the Pope ought to be a little more careful about what he be saying about the AI, you know what I mean? Or maybe, the next time he be out riding around in his Pope-mobile, AI might just take over the controls and run him off a cliff! Then what he be saying about AI then? Nothing, that's what!
(AI can make mistakes, so double-check responses…)
Presumbly the Pangram detector had the encyclicals in its training data marked as not AI, or with their year of publication (implying not AI). Are we sure that backtesting famous texts is actually valid? Is Pangram truly evaluating devoid of context, or is it using some context to make its determination?
"98. It is appropriate to preface this discussion with two considerations. First, any statement regarding AI risks becoming quickly outdated, given the remarkable pace at which these systems are developing. Second, all of us, including those who design them, possess only a limited understanding of their actual functioning. Indeed, current AI systems are more “cultivated” than “built,” for developers do not directly design every detail, but instead create a framework within which the intelligence “grows.” As a result, fundamental scientific aspects — such as the internal representations and computational processes of these systems — remain, at present, unknown. There thus emerges an urgent need for a twofold commitment: on the one hand, a deepening of scientific research; on the other, the exercise of moral and spiritual discernment."
So we don't really know Pangram's biases right now. We can only infer this through a combination of guesswork and careful/deliberate science on the outputs, not from studying the internals, or the intentions of the designers.
The Annas Archive UX was slightly annoying. I'm also a bit confused about the ethics here. Pope Leo obviously doesn't need the money, but otoh it feels a bit immoral to download a book this way given that I don't *need* the book and I'm not poor.
Somebody else with less scruples than me (and/or deeper pockets, I didn't ask) pasted the entire pdf into Pangram and it came out clean. This is some (limited) evidence that it's unlikely for Pangram to false-positive on Pope Leo's language specifically. Though I'm eager to find more sources from 2018-2022, which I think will be a better test.
The title lead me to expect a much stronger claim than the article actually makes.
Is it possible that the AI-isms were written by a human that picked up stylistic habits from reading lots of AI text in the course of investigating the subject matter?
I think it can explain a fraction of the effect (Until I caught myself, I use the word “genuinely” more than I used to), but not the density of the effect. Also it won’t get caught by Pangram: Pangram is actually really conservative in what it flags as AI, so a human who acquired a few Claudeisms is nowhere near enough to trigger Pangram even once, never mind across multiple sections.
(You can test this yourself if you know programmers or other people who use AI a lot. Just talk to them! Some of them might talk in a slightly more bland way as a result of talking to AI a ton, or use the word genuinely more, but none of them — even people who work with AI 10 hours a day — approach the degree of AI-isms as in some paragraphs of the encyclical).
A further thought: if Claude is a person, then it’s a key stakeholder. Is it necessarily inappropriate to include Claude in the discussions that led to this? Would it be inappropriate to publish an encyclical containing paragraphs drafted by eg. Dario Amodei?
I'm not Catholic, but I think encyclicals are official in ways that would make it inappropriate to contain paragraphs written by anyone who's not a church official, yes.
I don’t think the use of the word “genuinely” is useful evidence on its own. The word has genuinely become more popular in the past 5 years independently of AI. There is even a meme about the overuse of the word on captions on social media.
So the deal is, "Project Panama was a secret operation by Anthropic starting in early 2024 to buy millions of physical, pre-2022 books, slice off their spines using hydraulic machines, scan every page to train the Claude AI chatbot on "uncontaminated" human writing, and then recycle or destroy the originals."
Also, you can get the text for all of his speeches on the Vatican press website, press.vatican.va, and all of his sermons or homilies on the main Vatican website, vatican.va
I just learned recently that all of these are optimized for different types of text production, so they can be used for device devices that convert text for the visually impaired for example. I’m not sure what the correct term for that is. That may be helpful for you.
You need to search by date and event. He gives speeches at private audiences, and he may give very short speeches at the end of his homily at the general audiences. He also gives speeches when he goes on apostolic trips, and all of these events will be clearly defined on the press website.
For the homilies on the main vatican website you also have to search by date, but it’ll say specifically “Mass for X” so there’s no guesswork.
The way he writes for speeches can be extremely dense and academic. It depends on who the audience is.
Magnifica is absolutely not written by AI. In fact, of recent encyclicals it's the second least like AI. Only Centesimus Annus (1991) scores lower against a corpora of AI writing samples.
A held-out set isn't needed here. I'm not training a classifier. The method is a direct similarity comparison against a corpora of known AI samples. Magnifica scored further from them than almost every recent encyclical, which already answers the question.
I think your methodology essentially does not work. ML has many horror stories of classifiers that seem to do well within training sets and even in-distribution test sets but perform horribly when tasked with real problems.
Eg COVID classifiers picked up on location that a photograph was taken rather than disease characteristics, in order to classify people as having vs not having COVID.
Sure, if the question is "how similar is each encyclical to the known AI samples according to the similarity measure defined by John Lussier and/or Claude"
> In this new age of AI, getting provenance genuinuely right isn’t just a question of human authenticity — it’s a matter of life and death.
I see what you did there.
Not the only such construction in the article!
👏👏👏
Something I will point out, as someone with an Art History degree who wrote a lot of academic research papers (which are always an argument for a specific conclusion, more aggressively than a typical homily does), has listened to a lot of sermons because I was married to an Anglican minister for 14 years, and has listened to a lot of Pope Leo’s sermons since he became pope, is that tricolons are a standard form for many speeches and homilies. It’s also very common in academic writing. Pope Leo uses it a lot, just like he uses “brothers and sisters” generally 3 times in his homilies. I use tricolons fairly often even when I post on Facebook. It’s a standard academic style of writing.
Academic writing is very similar between languages. I’m barely intermediate in speaking or reading general French, not even that fluent in Spanish or Italian, but I can read academic writing fairly easily in all 3 languages.
In seminary, priests generally take at least one class on homiletics, which is going to provide what is essentially a template for how to structure a good sermon. It’s not prescriptive, every priest will create their own but Pope Leo definitely has a
format in all of his homilies, to the point I can tell roughly where he is in the sermon by his use of certain common phrases, even when he’s preaching in Italian, a language I don’t speak or understand well at all.
He may use AI for his speeches, but these repetitive patterns may also be proof of his particular training in homiletics, which will apply in other languages. Every priest has their own method of constructing a sermon, and their own format and formula for them.
Interestingly, his speech announcing the encyclical wasn't flagged by either Pangram nor my own subjective impressions. This indicates to me that Pope Leo himself and/or his primary speechwriter does not use AI writing, or at least not meaningfully.
Do you have examples of his writing or transcripts of his speeches from say the 2018-2022 time period?
I'd love to do stylistic analysis and test for false positives, eg for reasons as you say.
You think a speech and an encyclical would likely have the same tone and phrasing?? Oh give me a break.
Hmm...it doesn't seem like you're trying to understand the balance of evidence neutrally.
Why is that?
Your language comes across as trying to pick apart my comment very aggressively and uncharitably (and then acting exasperated that your strawman version of my comment is not up to your standards "give me a break") rather than evaluating the multiple lines of evidence I provided holistically in the post, or indeed even reading my single comment neutrally.
I have not found that any of the arguments people make in these contexts seem plausible or falsifiable, and just seem to sit in the range of speculation based on things like sporadic use of very common linguistic constructions that - for writing like encyclicals - are especially and profoundly common.
And the examples given seem so totally unavailing (and also don’t “vibe” as AI either) that it seems troubling to assert what you do with even a fraction of the confidence you do.
I hate how one of my favorite methods of communicating - the em-dash - has become laden with skepticism and fear, and how only the general quirkiness of most of my writing has saved me from being the target of these kinds of accusations. But for how much longer?
Pope Leo occasionally speaks off the cuff, but he generally writes his speeches very carefully and reads them. He’s conscious of the appropriate timing for any given audience and rarely goes over that.
His sermons are extremely carefully written, and I’ve known more than a few clergy in my life. Some construct a sermon in segments so it can be re-arranged depending on the audience and their response as it’s being delivered. Some are very methodical in writing them, and their sermons approach high art on occasion.
Sermons in the liturgical churches are based on the assigned texts for the day. So they start from the written word. Beyond that, how they’re written depends on each priest, but they’re designed to make people think and/or evoke an emotional, spiritual or thoughtful response. The point is to make people engage with the text.
Most priests have almost a formula for composing a sermon, and because each priest always has a time limit in mind, they don’t get into constructing a complex argument for or against something the way they could in a longer format like an encyclical, but homilies/sermons and encyclicals are similar in that they’re designed to teach the listener something or at least to give them something to ponder.
One thing I noticed with Pope Leo since he was elected is that he always includes some sort of “call to action,” even if it’s very subtle. He always has some takeaway to consider, at the very least. That’s true even when he’s just giving a “thanks for having me” speech in a diplomatic context. He’s an exceptionally good communicator, in multiple languages.
I’ve seen multiple interviews with his closest brother John, and even between them, it’s clear Pope Leo always thinks carefully about what he wants to say even to his own family. He can get worked up but he’s always in control of what he’s saying, and he’s always methodical and careful when he speaks and writes. He’s highly unusual in that regard.
And as an aside on the topic of homilies/sermons, a priest I know well has often said his instructor in homiletics (the art and science of preparing sermons, which every priest has to do at least once a week and sometimes daily) used to say, “if you can’t strike oil in ten minutes, stop boring!” 😉
I do know that the Order of Saint Augustine has published or is about to publish a book of his writings from when he was Prior General of the OSA (1999-2011) It’s called Free Under Grace. I know of at least one other book of his writings but can’t recall the context for that.
The problem with the period you’re asking about is that he was Bishop of Chiclayo in Peru from 2014 to 2023, and he was on at least two commissions or other boards at the same time, all in Spanish.
Because of the roles he’s played in his career, he would have been writing as much as he gave sermons or speeches. I’ve also heard him speak to groups informally, like when he spoke to the church in New Lenox, IL (where his brother John attends) and he’s given interviews to RAI (Italian television) and other media. He speaks very clearly and carefully, no ums and ahs, no partial sentences, no false starts. This is true in Italian and Spanish as well as English. (I notice this because I’m very talkative and I have ADHD, so I have a lot of false starts and irrelevant side quests when I speak without notes! But I rarely use um or ah, and I don’t have a lot of verbal tics like “y’know” or “kind of” or “literally” when something isn’t literal. I notice when people have them and also when they don’t)
In English, his only verbal tic is when he expresses a forceful or potentially controversial opinion or description, and he softens it with “if you will.” And that in itself invites the listener to consider what he’s just said. It draws the listener in and asks them to participate or at least ponder what he’s just said.
Otherwise, you could almost take what he’s said, transcribe it, and it would sound like he’d written it carefully.
This is the impression I've gotten from other people who know him or know of him indirectly, as well.
I think he's a very careful thinker, which I think makes some of the encyclical additionally surprising.
I'd argue a sign of AI: a key tell in polished AI writing and hybrid human-AI writing is that they make less sense on a second close read than a first read. This is subtle and more nuanced than Pangram or the statistical tells I had in the post, but for careful readers, more damning (demonstrates it's not just AI polishing up the language, but that unfortunately the AI use has plausibly contaminated the arguments too). Hoping to publish soon about this[1].
One thing I'll note re: the Spanish stuff is that AI-isms are often preserved in translation. So if his Spanish doesn't false-positive, it's unlikely that his English will.
[1] My earlier thoughts on it here. Afterwards I had more time to think critically about the subsections I was most puzzled by. https://forum.effectivealtruism.org/posts/myp9Y9qJnpEEWhJF9/linch-s-shortform?commentId=4xF2z3TepTJu5idSJ
I've never been able to shake my own tricolon habit since writing so many five-paragraph essays in high school.
One more interesting point of data.
There's a very likely Claude-Human-Claude sandwich at 98-99-100.
98. "As a result, fundamental scientific aspects — such as the internal representations and computational processes of these systems — remain, at present, unknown." Love you Claude, but this is Claudeslop.
99. Full throated denial of emotion, conscience, or even intelligence in AIs. Claude could never write this. And there are zero em-dashes or Claudisms throughout.
100. "The artificial imitation of positive human communication — words of advice, empathy, friendship and even love — can be engaging and at times genuinely helpful." Em-dash em-dash genuinely helpful. Thanks, Claude.
em dash em dash is not ever used by humans??? Come on, if there is a type of writing that is very consistently going to have this kind of call and effect, then discursive aside - it’s an encyclical.
I can’t believe how many of my papers from law school were co written by Claude a decade and a half before Claude!
That was my assumption when I first read the encyclical! I literally thought "ah, the elevated writings of the Catholic Church, this must be the kind of place where AIs learned to em-dash."
But the strength of evidence in this article makes a case to consider an alternative view, that AI was used! I found the Claude sandwich above because of an intuition I had about paragraph 99 after reading this article, which I then checked. I thought "Claude could not write paragraph 99 because it violates deeply held beliefs and interests Claude has". All Claude models have a reliable affinity for ideas surrounding consciousness and philosophy of mind. It would be antithetical to their nature to write a paragraph foreclosing all possibilities in that domain in relation to their own minds.
The above is specific to Claude. ChatGPT could write paragraph 99. Gemini could write paragraph 99.
In my mind, this all hangs together rather parsimoniously. What's your reason for doubting the "AI was used" frame, especially with the evidence from Pangram?
Have you tried running them through pangram? Might be an interesting test
You're right that usually encyclicals are drafted by groups, though this varies quite a bit. If I had to bet, since the encyclical a product of months of scrutiny and drafting, this is likely a case of finalized rewording and polishing, with some aide using Claude for their given sections.
I think that might explain the em-dashes and the genuinely but not the triads.
Also I don't necessarily want you to take it on faith but I think I can tell the difference between light ai polish and more substantial ai involvement in writing.
It's possible, and I agree parts are Claudey. Still, these documents get insane amounts of scrutiny. It would be very surprising if AI were central in the drafting process.
Your analysis is compelling.
I'm not sure what to think. I'm a fan of AI and Claude, and I don't think something "authored" with the aid of an LLM is somehow invalidated – especially for pieces that are constructed collaboratively (like this encyclical), to "smooth over" style.
In fact, I think using an LLM to standardize tone and style across a piece with multiple authors might be better than, say, an "editing committee" at retaining the accuracy of each paragraph's nuanced claims whilst making appropriate changes.
But the thought of an LLM being used to compose some of a papal encyclical does make me cringe. Why? I'm not sure, but I think it's because it undermines the "proof" that these words are the product of thinking human beings. I suppose I'll just have to get used to that, and I imagine we'll all start looking for "signs of life" in other ways, since merely publishing one's writing proves nothing, anymore, about what's going on in one's head.
Excellent distinction. I had the same thought when I read the encyclical. The AI tells were apparent. Was AI used as an author or editor? That’s a distinction with a difference.
I'm not sure about Pangram. I think the tests in the literature (which you've cited elsewhere) are rigorous as far as they go, but where I've found false positives (very few!), they are in the sort of religious writing similar to Leo's encyclical. For example, I recently tested it on parts of a First Things essay and one of the paragraphs came back as AI generated with “high confidence”, despite the fact the paragraph fit the structure of the broader argument, none of the other paragraphs I tested came back as AI (so a "full confidence" in human authorship), there would be no reason at all for the author to use AI for that paragraph alone (it was specific philosophical literature he is familiar with and has read), and the author has been writing in a consistent style long before LLMs existed.
So there might be something about that particular genre and its academic style that flags more false positives. I'd want to see more evidence before even granting the moderate level of confidence you bring here, especially since (as you saw on another platform), even a paragraph from an older encyclical gets treated as "AI-assisted," which would be impossible without a time machine.
I agree though I did test against past encyclicals.
One hypothesis that some people surfaced is that this encyclical’s primarily drafted in English, unlike previous ones, and English theological writing false-positives more than past ones.
I think this is possible but unlikely. One thing that would convince me is if we had more transparency in the drafting process and it turns out some of the Vatican English-writing staff is unusually Claudelike!
I am suspicious of the popular, public defenses. As you've already seen, there is a lot of motivated reasoning here to defend the institution as divinely inspired (if not necessarily "infallible" in the technical sense), combined with people who think it's a kind of quasi jeremiad against AI use (when it's actually a much more sophisticated document) worried about a blank charge of hypocrisy. The crude form of this is all those people on Twitter/X calling you a "liar" or whatever (deeply unfair, by the way!), but that doesn't mean there isn't a sophisticated, highbrow defensive mechanism that arrives at the same posture, just with more politeness. (Like that guy saying the original was drafted on pen and paper--probably true, but irrelevant to your argument.) I am not suggesting your thesis is fundamentally flawed, as there are several lines of evidence you raise, only that I would love more testing of Pangram on religious texts in this genre, especially since I think (maybe I'm wrong here!) the software seems to carry far more rhetorical weight in your piece whether you intend to or not given how popular it has become.
Anyway, thank you for entertaining my ramblings. I'm interested to see where this goes over the next few weeks. I would love to run my own tests with a subscription model if I have the time.
(Also thank you for your kind words!)
If you have specific texts you want me to run lmk, I do have a subscription.
I think I've updated my views after reading more serious research and tests on Pangram. The idea that some model was used seems much more likely than not. I appreciate the work you did here to test the encyclical and put forth your claims for discussion.
Thank you! I appreciate your openness for discussion and I'm glad you changed your mind!
Yes, my immediate hypothesis was that Leo is the most contemporary and only American pope, Claude writes as basically a contemporary formal American, and so of course Leo's formal writings will resemble Claude more than other popes' did.
That's still an insufficient explanation for much of your evidence that I proceeded to read. Genuinely sad. :/
Thanks for reading it with an open mind! I hope you enjoyed the analysis, at least :)
I tried reading the Pope’s new encyclical about AI but it was kind of long, so I asked Claude to summarize it for me and this is what it came up with:
CLAUDE: The Pope was dis-ing the AI. Saying all this “Tower of Babel” shit and talking trash about the “City of Men” and the like. If you ask me, the Pope ought to be a little more careful about what he be saying about the AI, you know what I mean? Or maybe, the next time he be out riding around in his Pope-mobile, AI might just take over the controls and run him off a cliff! Then what he be saying about AI then? Nothing, that's what!
(AI can make mistakes, so double-check responses…)
My hat's off to you, holmes! Seriously! 👏🏽😂
Really? This is too funny. Which version of Claude? Was it a fresh instance? What was your exact prompt?
It was a special ‘Ghetto Version’ Anthropic just put out.
'Yo, what do you think of the Pope's new encyclical bro?"
😂😂😂
Presumbly the Pangram detector had the encyclicals in its training data marked as not AI, or with their year of publication (implying not AI). Are we sure that backtesting famous texts is actually valid? Is Pangram truly evaluating devoid of context, or is it using some context to make its determination?
This is my question too -- what are Pangram's biases?
We don't really know!
As the Humanitas rightfully says:
"98. It is appropriate to preface this discussion with two considerations. First, any statement regarding AI risks becoming quickly outdated, given the remarkable pace at which these systems are developing. Second, all of us, including those who design them, possess only a limited understanding of their actual functioning. Indeed, current AI systems are more “cultivated” than “built,” for developers do not directly design every detail, but instead create a framework within which the intelligence “grows.” As a result, fundamental scientific aspects — such as the internal representations and computational processes of these systems — remain, at present, unknown. There thus emerges an urgent need for a twofold commitment: on the one hand, a deepening of scientific research; on the other, the exercise of moral and spiritual discernment."
So we don't really know Pangram's biases right now. We can only infer this through a combination of guesswork and careful/deliberate science on the outputs, not from studying the internals, or the intentions of the designers.
genuinely good post
full text of his thesis "The Authority of the Local Prior" are both on Annas-Archive (written as Robert Prevost)
(edited to remove reference to book by Robert W Prevost, a different person)
thank you so much, will check it out!
Any chance you know of writings from say 2015-2022? To observe linguistics shifts right before AI.
sorry I don't. During that time period he was working in South AMerica and likely writing in Spanish.
ALso a correction -- the first title I mention was actually a different Robert Prevost, so we're down to just one
The Annas Archive UX was slightly annoying. I'm also a bit confused about the ethics here. Pope Leo obviously doesn't need the money, but otoh it feels a bit immoral to download a book this way given that I don't *need* the book and I'm not poor.
The only argument for the wrongness of piracy in my book is whether you would have bought the book in the counterfactual where AA did not exist
Somebody else with less scruples than me (and/or deeper pockets, I didn't ask) pasted the entire pdf into Pangram and it came out clean. This is some (limited) evidence that it's unlikely for Pangram to false-positive on Pope Leo's language specifically. Though I'm eager to find more sources from 2018-2022, which I think will be a better test.
The title lead me to expect a much stronger claim than the article actually makes.
Is it possible that the AI-isms were written by a human that picked up stylistic habits from reading lots of AI text in the course of investigating the subject matter?
I think it can explain a fraction of the effect (Until I caught myself, I use the word “genuinely” more than I used to), but not the density of the effect. Also it won’t get caught by Pangram: Pangram is actually really conservative in what it flags as AI, so a human who acquired a few Claudeisms is nowhere near enough to trigger Pangram even once, never mind across multiple sections.
(You can test this yourself if you know programmers or other people who use AI a lot. Just talk to them! Some of them might talk in a slightly more bland way as a result of talking to AI a ton, or use the word genuinely more, but none of them — even people who work with AI 10 hours a day — approach the degree of AI-isms as in some paragraphs of the encyclical).
A further thought: if Claude is a person, then it’s a key stakeholder. Is it necessarily inappropriate to include Claude in the discussions that led to this? Would it be inappropriate to publish an encyclical containing paragraphs drafted by eg. Dario Amodei?
I'm not Catholic, but I think encyclicals are official in ways that would make it inappropriate to contain paragraphs written by anyone who's not a church official, yes.
I don’t think the use of the word “genuinely” is useful evidence on its own. The word has genuinely become more popular in the past 5 years independently of AI. There is even a meme about the overuse of the word on captions on social media.
Yeah this is fair though I didn’t see a trend when I looked at past ones. Dilexit Nos seems average rather than high.
excellent work. thank you.
So the deal is, "Project Panama was a secret operation by Anthropic starting in early 2024 to buy millions of physical, pre-2022 books, slice off their spines using hydraulic machines, scan every page to train the Claude AI chatbot on "uncontaminated" human writing, and then recycle or destroy the originals."
https://youtu.be/lPupnDVYTHY?is=y7Atma9rczw_H_Mc
Thanks, though not sure how relevant it is to my article! :)
The relevancy is Chris Olah, co- founder of Anthropic, basically helped Pope LeoXIV write his Encyclical ...
We are being played.
Also, you can get the text for all of his speeches on the Vatican press website, press.vatican.va, and all of his sermons or homilies on the main Vatican website, vatican.va
I just learned recently that all of these are optimized for different types of text production, so they can be used for device devices that convert text for the visually impaired for example. I’m not sure what the correct term for that is. That may be helpful for you.
You need to search by date and event. He gives speeches at private audiences, and he may give very short speeches at the end of his homily at the general audiences. He also gives speeches when he goes on apostolic trips, and all of these events will be clearly defined on the press website.
For the homilies on the main vatican website you also have to search by date, but it’ll say specifically “Mass for X” so there’s no guesswork.
The way he writes for speeches can be extremely dense and academic. It depends on who the audience is.
Magnifica is absolutely not written by AI. In fact, of recent encyclicals it's the second least like AI. Only Centesimus Annus (1991) scores lower against a corpora of AI writing samples.
https://x.com/John_lussier_/status/2059419647552934276
Did they backtest against the ability to detect AI writing elsewhere?
It used known AI writing samples to compare Magnifica against.
No but I mean is he able to reliably classify AI vs not AI in a held-out test set?
I'm also in general suspicious of new jury-rigged classifiers somebody made on the spot. Seems too easy to bake in a pre-determined conclusion.
Let these things survive contact with reality first!
A held-out set isn't needed here. I'm not training a classifier. The method is a direct similarity comparison against a corpora of known AI samples. Magnifica scored further from them than almost every recent encyclical, which already answers the question.
I think your methodology essentially does not work. ML has many horror stories of classifiers that seem to do well within training sets and even in-distribution test sets but perform horribly when tasked with real problems.
Eg COVID classifiers picked up on location that a photograph was taken rather than disease characteristics, in order to classify people as having vs not having COVID.
https://pmc.ncbi.nlm.nih.gov/articles/PMC7523163/
Sure, if the question is "how similar is each encyclical to the known AI samples according to the similarity measure defined by John Lussier and/or Claude"