apply() guarded on mic.getProcessor(), which is only set after setProcessor()
resolves (after the whole wasm/model load), so mic-published followed by
camera-published constructed two processors — two AudioContexts, two model
loads, double lock hold time. Track the in-flight processor per room, skip
apply() while one is pending, and destroy a pending processor if the module
is torn down before setProcessor resolves. Unit-tested.
Fixes#10
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PPmy3tPq869XDW4njjVaKA
The in-source ML denoiser is gated on lotusDenoiseSource (see lotusDenoise.ts
gate), not the legacy build-time shim flag lotusDenoise=ml. Correct the two
stale references (lotusDenoise.ts JSDoc + InCallView.tsx comment) to
lotusDenoiseSource=1.
Bump the embedded package to 0.20.1-lotus.2 for the next publish and update the
CI comment: the checked-in version now tracks the intended release, while a
pushed tag still wins (npm version "$TAG" overwrites at publish time).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Track-B audio-quality changes to reduce the "robotic/underwater" artifact.
- Dry/wet attenuation floor (default 0.15 ≈ -16 dB) blends a little of the raw
mic under the denoised signal so suppression can't fully collapse the noise
floor between words (the main cause of the RNNoise "underwater"/pumping
sound). Applied ONLY to the low-latency flat models (RNNoise/Speex); DTLN/DFN
add algorithmic latency that would comb-filter an undelayed dry mix, so they
rely on their own level instead. Tunable via `lotusDenoiseFloor`.
- Noise gate now runs AFTER the ML model, not before — gating the raw signal
fed hard-zeroed frames into the model and tuned the threshold on pre-denoise
levels.
- DeepFilterNet 3 noiseReductionLevel 80 -> 60: full strength was the main
"over-processed" contributor; 60 keeps voice natural.
Defaults are conservative and tunable; final values are meant to be dialed in
with real-call A/B listening.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Track-A robustness fixes from the engine review; no quality/model changes.
- H1: auto-resume the AudioContext on `statechange` if it suspends mid-call
(mobile backgrounding / audio interruption). Previously the dest node emitted
digital silence with no recovery — a silent mute of the sender.
- H2: `resumeCtx()` races `resume()` against a timeout. A suspended context can
only resume on a user gesture; the action can arrive via postMessage, so a
bare `await resume()` inside LiveKit's track-change lock could hang and
deadlock all later mute/unmute/device-switch. Now it proceeds and the H1
watcher heals it.
- M1: don't cache a REJECTED wasm fetch — a transient blip during a reconnect
used to permanently disable denoise for the session. Evict on failure.
- M2: activate denoise off `allConnections$` (local participant's connections)
instead of `livekitRoomItems$`, which excludes the local participant and only
surfaces rooms with a remote member — so denoise now also runs when you're
alone and no longer couples to a remote-render concern.
- Context lifecycle: `closeContext()` removes the state watcher before closing;
`ensureContext()` closes a half-initialised context on any failure (no leak).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The host already sets lotusDenoise=ml and injects its getUserMedia shim;
reusing that flag would double-process audio the moment this fork ships.
Gate the in-source engine on lotusDenoiseSource=1 instead, so the fork is
inert on deploy and the host cuts over explicitly (set the flag + drop the
shim) when ready.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Review found native dynamic import() of the DTLN/DeepFilterNet ESM
resolves "./denoise/…" against the bundled JS chunk's URL (-> /assets/…)
not the document, so those two models 404'd and silently fell back to raw
mic in the default config. Resolve the asset base to an absolute
same-origin href against the document; addModule()/fetch() accept absolute
too, so all three load paths stay consistent. (rnnoise/speex were
unaffected since addModule resolves against the document.)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Faithful port of cinny's proven pipeline into the TrackProcessor, closing
protocol gap F3 (host offers rnnoise/speex/dtln/deepfilternet; only the
first two existed in-source).
- Fix real bug: gate worklet registers as "noise-gate" (hyphenated), not
"noiseGate" — the gated path would have failed to construct the node.
- Per-model sample rate: DTLN runs at 16kHz, others 48kHz (worklets don't
resample); verify the context actually got the rate, else fall back.
- resume() a suspended context (host postMessage isn't a gesture).
- DTLN via dynamic-imported @workadventure helper (bypassUntilReady);
DeepFilterNet via dynamic-imported ESM + DeepFilterNet3Core pointed at
the self-hosted base. Same-origin base (kept from the C1 fix) makes these
dynamic imports safe.
- Prefer SIMD rnnoise.wasm with non-SIMD fallback; cache wasm per URL.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Denoise deep-review (CRITICAL): restart() read opts.audioContext, which
LiveKit does NOT pass on restart — so reconnect (the A7 scenario) and mic
device-switch threw after stopping the old track, leaving the mic SILENT
(A7 reintroduced). Fix:
- Processor owns a dedicated 48kHz AudioContext (sapphi worklets require
48kHz; H1), reused across restart, closed on destroy.
- restart() never throws and never leaves a stopped track on the sender:
builds the new graph first, then disposes the old; on failure degrades
to RAW mic audio rather than silence.
- Cache wasm per URL (no re-fetch each reconnect); gate threshold default
-45 and accept an explicit 0 (M2); document the cross-repo asset contract.
Protocol audit:
- Non-silent warning when an unsupported denoise model (dtln/deepfilternet)
is requested instead of silent rnnoise fallback (F3).
- Correct the call_state enum comment (immediate error-reply, not 10s) (F2).
Build/CI audit:
- Stamp VITE_APP_VERSION in CI; document the vX.Y.Z-lotus.N version scheme.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Holistic security audit findings:
- C1 (CRITICAL): force lotusDenoiseBase to same-origin before it reaches
audioWorklet.addModule()/fetch — a crafted call-link param could
otherwise load attacker JS/WASM as a worklet processing the live mic.
Non-same-origin/malformed values fall back to bundled ./denoise/.
- H1 (HIGH): gate audio-inject behind explicit lotusAudioInject=1 (still
acks the action so no transport hang) — it publishes under the local
user's identity, so it must not be silently armed for every call.
- M1 (MED): cap the decoration roster at 512 entries.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Implements RNNoise/Speex noise suppression as a LiveKit audio
TrackProcessor attached to the local mic track, replacing the host's
build-time getUserMedia monkeypatch. Because EC re-attaches the processor
on every (re)publish (LocalTrackPublished), denoise now survives EC's
mid-call reconnect — the root cause of A7 "mic dead after reconnect".
Reuses the worklet/wasm assets already shipped under ./denoise/ (no new EC
dependency); model/gate configured via lotusDenoise/lotusModel/lotusGate
URL params. Additive: no-op unless lotusDenoise=ml.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>