Proposed XMPP Extension: Jingle Audio/Video Conferences
The XMPP Extensions Editor has received a proposal for a new XEP. Title: Jingle Audio/Video Conferences Abstract: This specification defines a way to hold multiparty conferences with an SFU via Jingle. URL: https://xmpp.org/extensions/inbox/av_conferences.html The Council will decide in the next two weeks whether to accept this proposal as an official XEP.
Hi, Thanks for proposing this specification. I still fail to see why it is necessary or a good idea to create a new Jingle session for every participant stream. I see a lot of downsides without any obvious upside. As I understood, the reason to do this is that some SFU implementations use one WebRTC session per participant. However, there is no rule that says that a WebRTC session corresponds to a Jingle session. It's perfectly valid to translate multiple WebRTC sessions to multiple contents in a single Jingle session. There's no way for a client to know if an incoming call is part of already running conference call or independent, except for comparing the from-attribute if the iq stanza of an otherwise unrelated Jingle session. As of now, a single device can (technically) have multiple independent Jingle sessions and calls, even multiple between the same JIDs. While there probably aren't a lot of usecases for that (except for file share + single call at the same time), it seems inappropriate to disallow this by giving two calls with the same entity a special meaning (aka saying they form a conference) which is effectively what this XEP does. When multiple sessions are used, each stream needs to negotiate an additional ICE (+DTLS) session, because BUNDLE grouping is only defined within a Jingle session, not across sessions. This means unnecessary delay when new participants join. I understand that some available SFUs don't BUNDLE different participants, but that doesn't mean it's always a bad idea to support that. The current design doesn't allow the SFU to mix audio signals. While mixing is not strictly a task for SFUs, hybrid SFUs that mix audio and forward video are not uncommon for large A/V setups, so supporting this also would be reasonable. COIN already does support to announce audio mixing, but that by design wouldn't work if each participant needs to get their own Jingle session. Marvin On Tue, 2024-08-06 at 15:06 +0000, Daniel Gultsch wrote:
The XMPP Extensions Editor has received a proposal for a new XEP.
Title: Jingle Audio/Video Conferences Abstract: This specification defines a way to hold multiparty conferences with an SFU via Jingle.
URL: https://xmpp.org/extensions/inbox/av_conferences.html
The Council will decide in the next two weeks whether to accept this proposal as an official XEP. _______________________________________________ Standards mailing list -- standards@xmpp.org To unsubscribe send an email to standards-leave@xmpp.org
+1 to what Marvin says. Best Regards, Sergei вт, 6 авг. 2024 г. в 19:52, Marvin W <xmpp@larma.de>:
Hi,
Thanks for proposing this specification.
I still fail to see why it is necessary or a good idea to create a new Jingle session for every participant stream. I see a lot of downsides without any obvious upside. As I understood, the reason to do this is that some SFU implementations use one WebRTC session per participant. However, there is no rule that says that a WebRTC session corresponds to a Jingle session. It's perfectly valid to translate multiple WebRTC sessions to multiple contents in a single Jingle session.
There's no way for a client to know if an incoming call is part of already running conference call or independent, except for comparing the from-attribute if the iq stanza of an otherwise unrelated Jingle session.
As of now, a single device can (technically) have multiple independent Jingle sessions and calls, even multiple between the same JIDs. While there probably aren't a lot of usecases for that (except for file share + single call at the same time), it seems inappropriate to disallow this by giving two calls with the same entity a special meaning (aka saying they form a conference) which is effectively what this XEP does.
When multiple sessions are used, each stream needs to negotiate an additional ICE (+DTLS) session, because BUNDLE grouping is only defined within a Jingle session, not across sessions. This means unnecessary delay when new participants join. I understand that some available SFUs don't BUNDLE different participants, but that doesn't mean it's always a bad idea to support that.
The current design doesn't allow the SFU to mix audio signals. While mixing is not strictly a task for SFUs, hybrid SFUs that mix audio and forward video are not uncommon for large A/V setups, so supporting this also would be reasonable. COIN already does support to announce audio mixing, but that by design wouldn't work if each participant needs to get their own Jingle session.
Marvin
On Tue, 2024-08-06 at 15:06 +0000, Daniel Gultsch wrote:
The XMPP Extensions Editor has received a proposal for a new XEP.
Title: Jingle Audio/Video Conferences Abstract: This specification defines a way to hold multiparty conferences with an SFU via Jingle.
URL: https://xmpp.org/extensions/inbox/av_conferences.html
The Council will decide in the next two weeks whether to accept this proposal as an official XEP. _______________________________________________ Standards mailing list -- standards@xmpp.org To unsubscribe send an email to standards-leave@xmpp.org
_______________________________________________ Standards mailing list -- standards@xmpp.org To unsubscribe send an email to standards-leave@xmpp.org
Le mardi 6 août 2024, 18:52:04 UTC+2 Marvin W a écrit :
Hi,
Thanks for proposing this specification.
I still fail to see why it is necessary or a good idea to create a new Jingle session for every participant stream. I see a lot of downsides without any obvious upside. As I understood, the reason to do this is that some SFU implementations use one WebRTC session per participant. However, there is no rule that says that a WebRTC session corresponds to a Jingle session. It's perfectly valid to translate multiple WebRTC sessions to multiple contents in a single Jingle session.
There's no way for a client to know if an incoming call is part of already running conference call or independent, except for comparing the from-attribute if the iq stanza of an otherwise unrelated Jingle session.
As of now, a single device can (technically) have multiple independent Jingle sessions and calls, even multiple between the same JIDs. While there probably aren't a lot of usecases for that (except for file share + single call at the same time), it seems inappropriate to disallow this by giving two calls with the same entity a special meaning (aka saying they form a conference) which is effectively what this XEP does.
When multiple sessions are used, each stream needs to negotiate an additional ICE (+DTLS) session, because BUNDLE grouping is only defined within a Jingle session, not across sessions. This means unnecessary delay when new participants join. I understand that some available SFUs don't BUNDLE different participants, but that doesn't mean it's always a bad idea to support that.
The current design doesn't allow the SFU to mix audio signals. While mixing is not strictly a task for SFUs, hybrid SFUs that mix audio and forward video are not uncommon for large A/V setups, so supporting this also would be reasonable. COIN already does support to announce audio mixing, but that by design wouldn't work if each participant needs to get their own Jingle session.
Marvin
Hi Marvin, thanks for the feedback. As we have discussed during last Berlin sprint, I've followed closely the implementation of the SFU I've used for the implementation (Galène) for simplicity, and to have something working quickly. I'm totally open and willing to move gradually to another design, and notably to use a single session. I'm still wondering if it would make sense to have both approaches or if having a single Jingle session is always the best way to go (also regarding the facility of implementation which is a design goal). The current design has the advantage to be very easy to implement, and relatively similar to what is done with MUJI (without the MUC part, which is not necessary here). The peer JID is indeed the way to differentiate the call: the disco identity is here to indicate that it's an A/V conference room (and XEP-0298 'isfocus' attribute is used during the Jingle negotiation as additional indicator). The SFU service must not do independent call, it's A/V conferences only, so there is not way to mix it. A file sharing can still happen with an SFU service/av conference room (it's just another session with jingle-ft instead of A/V streams). Yes the ICE negotiation indeed add delay, this is also noted on Galène. I think that this is acceptable as first implementation, and this can evolve to a better design with following revisions. While supporting audio mixing would be nice, I don't think that this is necessary for first implementation. This also can be a future addition when the design will be updated. The goal here is to have something very simple, straightforward to implement, so XMPP ecosystem can have several clients with multiparty conferences soon. I have already a SFU service and client implementation using this specification. Again I'm totally open to move this gradually to a better design. However I think that the current one is acceptable as starting point, and is very easy to implement. On another topic, several people have informed me that there was a copy/paste error with the "overview" section appearing twice, it would be nice to remove one of them but I don't want to modify the protoXEP before council review. I'll do it at next revision. Note that I'll be on vacation in coming weeks. I'll check from time to time my message, but I'll be very slow to answer. Best, Goffi
Dear all, It appears that this proposal has been vetoed. However, there is no information available on this list regarding the reason for this decision. I had to consult the council logs, and have encountered a similar issue with my previous proposal. Could singpolyma please provide a detailed explanation of why this work was vetoed? Additionally, what steps would be necessary to revise the proposal so that it can be approved? I understand that "it is incompatible with what's needed for most SFUs," but I believe this statement is incorrect. This XEP is an adaptation of the protocol used in the Galène SFU. Furthermore, I have already indicated my willingness to evolve the proposal towards a single Jingle session, though for the initial draft, I chose to make a direct adaptation of the Galène protocol. Thank you for considering these points. On another note, while I understand that both the editor and council have heavy workloads, it would be beneficial if the results of council votes were shared on this list, possibly accompanied by relevant log excerpts. Thank you once again. Best regards, Goffi
Le mardi 20 août 2024, 18:35:52 UTC+2 Stephen Paul Weber a écrit :
I have already indicated my willingness to evolve the proposal towards a single Jingle session
Once it uses a single Jingle session, what is even left for the XEP to say that isn't covered by existing Jingle XEPs?
Hi Singpolyma, Thank you for your answer. Just because a specification is small doesn't mean it's not useful. Currently, there is no way to handle an A/V conference room in XMPP. This protoXEP specifies how to join an A/V conference room, use metadata, implement the discovery mechanism, determine the identity to use, and configure the room. Its concise nature is deliberate, aimed at facilitating easy implementation, but it remains a necessary component for making A/V conference calls. Without it, we have currently no way to make SFU-based conference calls. Furthermore, I plan to extend this specification by adding functionality to retrieve joined users and notify new joiners, as the SFU may choose not to send all participants' streams. I'd also like to mention that I have a working implementation of this protoXEP, utilizing a server component based on Galène. Regards, Goffi
Le mercredi 21 août 2024, 17:17:47 UTC+2 Goffi a écrit :
Le mardi 20 août 2024, 18:35:52 UTC+2 Stephen Paul Weber a écrit :
I have already indicated my willingness to evolve the proposal towards a single Jingle session
Once it uses a single Jingle session, what is even left for the XEP to say that isn't covered by existing Jingle XEPs?
Hi Singpolyma,
Thank you for your answer.
Just because a specification is small doesn't mean it's not useful.
Currently, there is no way to handle an A/V conference room in XMPP. This protoXEP specifies how to join an A/V conference room, use metadata, implement the discovery mechanism, determine the identity to use, and configure the room. Its concise nature is deliberate, aimed at facilitating easy implementation, but it remains a necessary component for making A/V conference calls.
Without it, we have currently no way to make SFU-based conference calls.
Furthermore, I plan to extend this specification by adding functionality to retrieve joined users and notify new joiners, as the SFU may choose not to send all participants' streams.
I'd also like to mention that I have a working implementation of this protoXEP, utilizing a server component based on Galène.
Regards, Goffi
So, what are the steps to make an updated proposal acceptable?
So, what are the steps to make an updated proposal acceptable?
Since you've already signalled you are ok to make the proposal use a single jingle stream (or a flexible number of streams) instead, that seems like a good thing here. I'm not sure what will be left after that, but certainly it would remove my main objection.
Le lundi 26 août 2024, 15:52:34 UTC+2 Stephen Paul Weber a écrit :
So, what are the steps to make an updated proposal acceptable?
Since you've already signalled you are ok to make the proposal use a single jingle stream (or a flexible number of streams) instead, that seems like a good thing here. I'm not sure what will be left after that, but certainly it would remove my main objection.
Alright, thanks for the update. My primary goal is to create a specification that enables users to easily use my SFU component. I understand several XMPP client developers are interested in implementing this feature, and I believe the current specifications might be too generic or insufficient for a straightforward implementation. I'll also work on a method to advertise/ retrieve joined users. If others have alternative strategies or wish to collaborate on this or a similar specification, please reach out so we can avoid duplicating efforts. Currently, I'm quite busy and won't be able to submit an updated proposal within the next several weeks. Thanks. Regards, Goffi
participants (5)
-
Daniel Gultsch -
Goffi -
Marvin W -
Sergei Ilinykh -
Stephen Paul Weber