To make a general point, I'm not "against" AI in principle, and don't consider myself the kind or "AI tech bro" type either. It is a tool, and like many tools can be misused, and needs careful wielding. I've had much success with it as a tool, and also some mis-steps.
Would I be happy with a specification that was originally written in, say, Chinese and then blindly Google-translated into English? Probably not. Would I be happy with a specification written by a Chinese speaker who used Google Translate to help express themselves? Sure. Would I want either declared? Nope, because the risk of tainting the perfectly sensible use of the second would risk alienating otherwise good submissions.
As usual, I'd note that I don't write my XEPs using LLMs, because I happen to enjoy writing them and was lucky enough to be born British, so get to write English like what it's meant to be. But I also like languages, and so I'm somewhat aware of how hard it must be to write precise technical specifications in a foreign language. My "best" language other than English is probably French, and there is absolutely no way I could write technical French of that quality - I can barely read it - and if I had to do so for some reason I'd absolutely be leaning heavily on tooling like LLMs. It'd be painful to have that work tossed out or sneered at simply because of that.
Indeed, I'd suggest there may be an argument that it's not in line with the Code of Conduct, as it's hardly welcoming to those who - for no fault of their own - are not fluent English writers.
Low-quality contributions are not new. What's changed is that they no
longer look like low-quality contributions. Identifying slop takes an
unfair amount of cognitive effort compared to the effort required to
generate and submit it.
I don't think this is quite the right take.
I also don't think we have an active problem, yet, in the XSF.
But I do think it's entirely likely we'll run into an increase in the kind of submissions which are easy-to-produce AI slop. The IETF has certainly seen this, though in the interesting manner that much of the AI-generated Internet Drafts there are (anecdotally) entire suites of drafts, like "IPv8".
On the other hand, LLMs can also be used effectively to write specifications, I'm sure. As an example, I suspect (without, yet, trying) that generating comprehensive examples from the text I've written ought to be possible (and if not, would be illuminating as to what is missing from the text). Certainly I've seen LLM reviews of specifications and been impressed at what nits (and serious issues) they spot.
Great tools, and thus can be abused greatly.
I firmly believe we should ensure we handle the abuse, not the use.
Knowing that a XEP submission came in part or even wholly from LLM
output changes the way I review a document.
This, you see, is an example of my point at the beginning.
LLMs are extremely good at
generating sensible-looking documents. If a sensible-looking document
comes from an experienced community member, I am likely to spend 80%
of my time assessing the higher level aspects of their submission - I
am not likely to cross-reference every namespace or referenced
document to see if they hallucinated it or not. One can certainly
argue that we should give every document a complete thorough review
right out of the gate, checking and cross-referencing everything, but
that's more effort than many of us can commit to and has never been
the case really. LLM documents tend to be longer, and this makes such
work even harder.
"We" - for some value of we - absolutely should give every document a thorough review of this form at some point prior to at least Stable. Right?
I suspect part of your stance here is the way that Experimental has shifted in nature from being "rough draft" to "almost ready", and thus shifting that review back to the submission point, which means any influx of "AI slop" causes a huge amount of workload for Editor and Council alike.
Putting aside philosophical argument against LLMs on principle for the moment, this feels like the pragmatic problem, this feels like the actual problem that needs addressing. We need to ensure that if - or more likely when - we get an influx of AI slop our process can handle it without driving "us" (again, for some values of "us") to exhaustion.
I would quite like to fix this rather large gap between what we document as our process and what we actually appear to do, anyway.
Instead, I am suggesting for many reasons already presented in this
thread, that we simply ask people to declare whether and how they used
AI in their contributions. If a submission says "this is entirely my
own work, but I used an LLM to review it" and another submission says
"this is output directly from an LLM, I've put no effort into checking
any of its contents" then I'm obviously going to treat those documents
very differently as a reviewer.
I refer you to Ralph's XEP-0076 comment.
The fact is that many people will be entirely unaware of LLM involvement in their spec writing (grammar checkers are a great example you've noted - I have no idea if Google Translate uses LLMs now), but others will ignore this and submit a suite of 60 page XEPs generated by some LLM without any understanding of utility.
We'll catch both, post-facto. We will penalize both, as per policy. The damage has already been done by the abuser, and will mostly affect the innocent user.
If nothing else, it will allow us to get an idea of how AI tools are
being used in our community. And things *could* be fine... if in 12
months we look back and see a bunch of in-progress documents with AI
use declared, but the documents are actually in great shape, then
we'll at least know we were worrying about nothing when it comes to
XEP quality. But if we consistently see that the documents authored
using AI are below our usual standards, then we can have further
discussions in the future about what kind/amount of AI usage is
appropriate in contributions. But right now, we're largely in the
dark, and that's scary. It leads to people jumping up and down every
time they see an emdash.
I think focusing on declaring what tools have been used risks a different kind of witch hunt.
I think we risk policing use, but tacitly allowing abuse.
Dave.
--
(This email was grammar checked automatically, though I didn't accept any suggestions. I have absolutely no idea, as previously stated, whether this involved the use of LLMs, so I asked an LLM, and therefore I've definitely used an LLM in the generation of this email, irrespective of the answer. Dagnabbit.)