💬 New comment on ntgcalls#62 Whole interpreter deadlocks when the app does I/O beside a live call - PyEval_AcquireThread behind a webrtc lock
by @osyris
Thanks - built dev at 6d4602d on macOS 15.5 arm64 / CPython 3.12, loads and runs under py-tgcalls 3.0.0.dev5 (3.0.0b18+dev.6d4602d).
The diff does more than I asked for, and both halves look right to me. synchronized_callback now copies the callback out from under mutex_ and invokes it unlocked, so no native lock is held across PyEval_AcquireThread`; and `py::call_guard<py::gil_scoped_release>() on the bound methods means a Python thread blocked inside a native call is no longer sitting on the GIL. That removes both sides of the inversion rather than one, which is the difference between "this ordering is now safe" and "this ordering no longer exists".
The concrete thing it unblocks for us: after that wedge we had to stop claiming the ntgcalls logger back at INFO, because the LogSink thread entering the interpreter on every line was one of the three threads in that dump - and since importing pytgcalls raises that logger from NOTSET to CRITICAL, leaving it alone means the media stack goes silent. That is the logging that made the call failures readable in the first place, so we've been running blind to keep the deadlock away. With this in, it goes back on.
Same ask as on #61: the wedge took a live call plus real event-loop churn beside it, so I can't reproduce it synthetically. If you cut a beta with this in it I'll run it against real calls with the media log turned back up and report back. Happy to run our own build in production instead if you'd rather not cut one.
Reply to this message to post a comment on GitHub.
💬 New comment on ntgcalls#61 SIGSEGV on ntg-work removing an incoming audio channel during a conference - null deref at +0x58, three identical crashes
by @osyris
Thanks - built dev at 6d4602d on macOS 15.5 arm64 / CPython 3.12 and it links and loads fine next to py-tgcalls 3.0.0.dev5 (3.0.0b18+dev.6d4602d).
Reading the diff, your fix explains the crash better than my guess did. I was looking at the queued packet tasks, but the dangling reference was the other way round: the receive channel still held the raw sink and the frame transformer after ~IncomingAudioChannel dropped channel_. Unregistering both inside the same worker_thread_.BlockingCall before nulling it closes it properly - the worker thread is serialized, so anything already queued for that ssrc either ran before the unregistration or finds nothing registered.
One question while you have it in your head: the same commit also stops building the audio_level_and_speech lambda when set_audio_level_and_speech is empty. Do you think that one was reachable here, or is it a separate latent one you spotted on the way past? We don't ask for speaking detection, so if it was reachable it would be the likelier of the two - though an empty std::function should throw rather than fault at 0x58, which is why I'd lean on the sink.
What I can't do from here is the test that counts. This only ever fired on real conferences and only sometimes - three times in three days, one 27 seconds in and one 26 minutes in - so a synthetic run proves nothing. We deploy from PyPI wheels, so if you cut a beta with this in it I'll put it in front of real conference traffic and report back either way. If you'd rather not cut one just for this, say so and I'll run our own build in production instead.
Reply to this message to post a comment on GitHub.
💬 New comment on ntgcalls#62 Whole interpreter deadlocks when the app does I/O beside a live call - PyEval_AcquireThread behind a webrtc lock
by @Laky-64
Can you try please with the latest version?
https://pypi.org/project/ntgcalls/3.0.0b19/
Reply to this message to post a comment on GitHub.
💬 New comment on ntgcalls#61 SIGSEGV on ntg-work removing an incoming audio channel during a conference - null deref at +0x58, three identical crashes
by @Laky-64
Can you try please with the latest version?
https://pypi.org/project/ntgcalls/3.0.0b19/
Reply to this message to post a comment on GitHub.
💬 New comment on ntgcalls#61 SIGSEGV on ntg-work removing an incoming audio channel during a conference - null deref at +0x58, three identical crashes
by @osyris
Running b19 in production since today - our daemon reports PyTgCalls v3.0.0dev5 powered by NTgCalls v3.0.0b19+dev.c82e5b2 and it has been up and healthy since.
I can't call it fixed yet, and I don't want to say otherwise: this fired three times in three days and then not once in the two days before we upgraded, so a quiet week on b19 is not evidence on its own. What would be evidence is the same conference traffic that produced all three crashes, and that happens here most days. I'll come back either way - if it goes a couple of weeks with real conferences and no ntg-work fault I'll say so plainly, and if it faults again I'll have another .ips with the same offsets to hand you.
One small thing worth knowing, since it cost me a few minutes: the v3.0.0-beta19 tag points at 38104dd6, which is master's head from 2026-08-03 and predates both fixes. The wheel is fine - it was built from dev at c82e5b20, and ntgcalls.__version__ says 3.0.0b19+dev.c82e5b2 - but anyone who diffs the tag to check whether the fix shipped will conclude it didn't. Might be worth pointing the tag at what was actually built.
Reply to this message to post a comment on GitHub.
💬 New comment on ntgcalls#62 Whole interpreter deadlocks when the app does I/O beside a live call - PyEval_AcquireThread behind a webrtc lock
by @osyris
Running b19 in production since today (NTgCalls v3.0.0b19+dev.c82e5b2 under py-tgcalls 3.0.0dev5), and the first thing we did with it was put the media log back.
That is the part I can already report as a win. Since importing pytgcalls leaves the ntgcalls logger at CRITICAL, claiming it back is the only way to see anything the stack says - and it was claiming it that fed the LogSink thread into the interpreter on every line, which is how it ended up as one of the three threads in that dump. We had it turned off to stay away from the deadlock, which meant running blind through exactly the failures we most needed to read. It's back at INFO now, with the b19 pin as the thing that makes it safe.
The wedge itself I can't call fixed yet. It took a live call plus a real chunk of event-loop work beside it, and we only ever hit it twice - so silence for a few days proves nothing. What it needs is the same shape of traffic on the new build, which happens here most days, and I'll report back either way.
Reply to this message to post a comment on GitHub.
💬 New comment on ntgcalls#58 P2P call never reaches CONNECTED on 3.0.0b15 — CONNECTING → TIMEOUT after 10s while key exchange completes and relays are valid
by @osyris
Now that we're on b19, your call on the LogSink floor turned out to be right twice over. Claiming the logger first, exactly as you described, is what we run - and your #62 fix is what made it safe to run, because claiming it was precisely what fed the log thread into the interpreter. The floor stays where you put it, and the advice now works as written.
Five reports closed in a week across two libraries, three of them commits the same day. Thank you - that is not the usual experience of filing a bug, and it seemed worth saying on its own rather than only when we come back with the next one.
Reply to this message to post a comment on GitHub.
𝐒ιᴅᴅʜᴀɴᴛDm kru sir ? Hindi idhr allowed nhi h
@HeySiddhant [6521935712] warned (1 of 3).
Due to: Keep the chat in English
ㄥ卂ㄩ尺乇几 ッ - 劳伦I mean regarding the fix
I lost the actual file, but I tested it in different ways, and it looks good.