On Wed, 8 Jul 2026 at 18:08, MSavoritias via Standards <standards(a)xmpp.org>
wrote:
1. It seems to be a consensus in the list that XEPs
can be rejected if
they are "bad" in some way. The issue here is that its not clear what
"bad" means. This would help the editor (or council) a lot for example
to be able to point at "something" that is public and clear. This can of
course also help for the Code of Conduct and for the Experimental
process among other things.
I think your last suggestion actually handles the "definition of bad",
or
rather, defines a "definition of good enough" that I think - with maybe
some tweaks - is ideal.
IMHO a good baseline is:
- The author of said XEP *CAN* give the copyright to XSF
Right. I think (hope!) we have this already. I refer to this as the
copyright warranty.
- The author of said XEP *CAN* explain why and how in
a XEP
I'm less bothered about this, but I see the logic, and think it's covered
by Guus's PR (of Ralph's words).
If you want clarification, I mean that obviously it's sensible to
understand the spec you're submitted in all its details, but I'm not sure
we need a warranty to that effect. It equally obviously does no harm,
though, so I've no objections either!
- The author of said XEP *HAS* communicated with the
XSF (in one of our
rooms, mailing list, summit, etc.) or is vouched by one of the members
of XSF before making the XEP that their approach is at least desired and
may make some sense for initial experiments.
I think this needs tweaking but the essential concept - that other people
support the submission - is really important.
This does two things:
- It means that the kind of mindless AI slop that I think we're all rightly
concerned about never gets traction - not because it's AI, but because it's
mindless slop.
- It also means that the early stage workload is offloaded from Council.
Council only need to take notice of submissions that actually gain any kind
of traction in the community.
I would s/one of the members of XSF/participants in the Standards SIG/
because we've not selected our membership on the basis of technical
scrutiny, and it just feels like if a ProtoXEP is getting positive
discussion and engagement on the standards list it's probably good for
Experimental. Council's effort is then a judgement call on that, rather
than having to scrutinise the specification itself as much.
This solves: XEP "dumps" where we don't know why it is like this, who
wrote this, is this in good faith, is this wanted/needed etc., solves
the work of somebody submitting XEPs without knowing what it is written
in there and also covers liability. Note that the first is already
written (and imo that makes llms unable to be used already as openjdk
and other projects have stated)
Your parenthetical opinion is far from universal.
First, the situation with code is radically different to that for text, and
second, the copyrightability of AI output varies heavily by jurisidiction
and human effort involved. Here in the UK, for instance, we have existing
primary legislation (from 1986!) that says LLM output is copyrightable.
Other jurisdictions have case law concerning the extremes of AI output, so
we know that a XEP coming from a US citizen solely generated by a single
prompt of something like "Create a new XEP for something" would likely not
be copyrightable.
But, this doesn't matter, because by requiring the author to warrant they
can assign copyright, we push that liability onto the author. Hoorah!
2. Regarding what approach we can have to actually
write this down some
examples are that *DO NOT* ban LLMs:
All these are code based, and our concern here is prose. We're protected in
any case because of the warranty we demand from submitters.
You can stop reading here if you want, the rest is just saying "I
understand the argument but reject it".
There's an argument that because the training data for LLMs was (possibly
misused) copyright code, the output is a derived work (in copyright terms)
of the original training data. This is supported because at least in some
cases in the early days, if you asked for particularly niche code, you'll
end up with output looking very similar to pre-existing projects. I have
not duplicated these results, and I absolutely tried hard - I have the only
open source code based on certain specifications which are not public, so
niche in some interesting ways, but I've seen other claims and have no
reason to doubt it can (or could) happen.
The problem with this is that few people are making this argument with
prose - people actually complain that AI generated prose doesn't look human
enough - and also nobody makes this argument with human generated anything.
I can fully assure you that this email is not considered a derived work of
Asimov or Bujold or Scalzi, despite the fact I have read all three an
almost unhealthy amount. Similarly, the code I write has never been
considered a derived work even though I certainly read other people's
copyright code to learn how to write it.
Whether or not the LLMs were originally trained with copyright works in
contravention of their licences is a whole other matter, but from a purely
legal perspective I think that's between the copyright owners and the LLM
developers. You are welcome to take an ethical stance on that one on a
personal basis.
Dave.