Remove the rebase hazard from CallViewModel: upstream's spotlightSpeaker$
auto-selection had been renamed to autoSpotlightSpeaker$ and an inline
manual-override + screenshare-coexistence block was spliced into
spotlightAndPip$. Both diverge from upstream and would conflict on every
rebase.
Restore spotlightSpeaker$ to its byte-for-byte upstream form and move the
[lotus #4] override into a pure wrapper, overrideSpotlight$(), invoked at a
single call point in spotlightAndPip$. Behaviour is unchanged: identical to
upstream while manualSpotlightUserId$ is null (the default), and preserves the
"pin a participant" and "focus camera during screenshare" (#4 / A5) rules when
the host sends io.lotus.focus_participant.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
playInjectedClip now stops any in-flight clip (via the existing idempotent
cleanup) before starting a new one, so rapid taps replace rather than
overlap/stack — no track leak. The host also debounces the button.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
AEC/AGC audit fix + two hardening items from the engine review.
- Add an `autoGainControl` capture param (UrlParams -> CallViewModel ->
ConnectionFactory audioCaptureDefaults), mirroring echoCancellation/
noiseSuppression. Defaults true (unchanged); the host sets it false only for
the ML tier so the browser's auto gain control doesn't fight the in-source ML
denoiser (pumping). Echo cancellation stays on. Tests cover the URL parse and
the audioCaptureDefaults wiring.
- L1: init() now closes the owned AudioContext on a build failure (was orphaned;
browsers cap live contexts, so repeated failures could exhaust them).
- L2: buildGraph() disposes its partially-built nodes on failure (disposeGraph
previously only cleaned the prior graph).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Track-B audio-quality changes to reduce the "robotic/underwater" artifact.
- Dry/wet attenuation floor (default 0.15 ≈ -16 dB) blends a little of the raw
mic under the denoised signal so suppression can't fully collapse the noise
floor between words (the main cause of the RNNoise "underwater"/pumping
sound). Applied ONLY to the low-latency flat models (RNNoise/Speex); DTLN/DFN
add algorithmic latency that would comb-filter an undelayed dry mix, so they
rely on their own level instead. Tunable via `lotusDenoiseFloor`.
- Noise gate now runs AFTER the ML model, not before — gating the raw signal
fed hard-zeroed frames into the model and tuned the threshold on pre-denoise
levels.
- DeepFilterNet 3 noiseReductionLevel 80 -> 60: full strength was the main
"over-processed" contributor; 60 keeps voice natural.
Defaults are conservative and tunable; final values are meant to be dialed in
with real-call A/B listening.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Track-A robustness fixes from the engine review; no quality/model changes.
- H1: auto-resume the AudioContext on `statechange` if it suspends mid-call
(mobile backgrounding / audio interruption). Previously the dest node emitted
digital silence with no recovery — a silent mute of the sender.
- H2: `resumeCtx()` races `resume()` against a timeout. A suspended context can
only resume on a user gesture; the action can arrive via postMessage, so a
bare `await resume()` inside LiveKit's track-change lock could hang and
deadlock all later mute/unmute/device-switch. Now it proceeds and the H1
watcher heals it.
- M1: don't cache a REJECTED wasm fetch — a transient blip during a reconnect
used to permanently disable denoise for the session. Evict on failure.
- M2: activate denoise off `allConnections$` (local participant's connections)
instead of `livekitRoomItems$`, which excludes the local participant and only
surfaces rooms with a remote member — so denoise now also runs when you're
alone and no longer couples to a remote-render concern.
- Context lifecycle: `closeContext()` removes the state watcher before closing;
`ensureContext()` closes a half-initialised context on any failure (no leak).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The host already sets lotusDenoise=ml and injects its getUserMedia shim;
reusing that flag would double-process audio the moment this fork ships.
Gate the in-source engine on lotusDenoiseSource=1 instead, so the fork is
inert on deploy and the host cuts over explicitly (set the flag + drop the
shim) when ready.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Review found native dynamic import() of the DTLN/DeepFilterNet ESM
resolves "./denoise/…" against the bundled JS chunk's URL (-> /assets/…)
not the document, so those two models 404'd and silently fell back to raw
mic in the default config. Resolve the asset base to an absolute
same-origin href against the document; addModule()/fetch() accept absolute
too, so all three load paths stay consistent. (rnnoise/speex were
unaffected since addModule resolves against the document.)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Faithful port of cinny's proven pipeline into the TrackProcessor, closing
protocol gap F3 (host offers rnnoise/speex/dtln/deepfilternet; only the
first two existed in-source).
- Fix real bug: gate worklet registers as "noise-gate" (hyphenated), not
"noiseGate" — the gated path would have failed to construct the node.
- Per-model sample rate: DTLN runs at 16kHz, others 48kHz (worklets don't
resample); verify the context actually got the rate, else fall back.
- resume() a suspended context (host postMessage isn't a gesture).
- DTLN via dynamic-imported @workadventure helper (bypassUntilReady);
DeepFilterNet via dynamic-imported ESM + DeepFilterNet3Core pointed at
the self-hosted base. Same-origin base (kept from the C1 fix) makes these
dynamic imports safe.
- Prefer SIMD rnnoise.wasm with non-SIMD fallback; cache wasm per URL.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Denoise deep-review (CRITICAL): restart() read opts.audioContext, which
LiveKit does NOT pass on restart — so reconnect (the A7 scenario) and mic
device-switch threw after stopping the old track, leaving the mic SILENT
(A7 reintroduced). Fix:
- Processor owns a dedicated 48kHz AudioContext (sapphi worklets require
48kHz; H1), reused across restart, closed on destroy.
- restart() never throws and never leaves a stopped track on the sender:
builds the new graph first, then disposes the old; on failure degrades
to RAW mic audio rather than silence.
- Cache wasm per URL (no re-fetch each reconnect); gate threshold default
-45 and accept an explicit 0 (M2); document the cross-repo asset contract.
Protocol audit:
- Non-silent warning when an unsupported denoise model (dtln/deepfilternet)
is requested instead of silent rnnoise fallback (F3).
- Correct the call_state enum comment (immediate error-reply, not 10s) (F2).
Build/CI audit:
- Stamp VITE_APP_VERSION in CI; document the vX.Y.Z-lotus.N version scheme.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Holistic security audit findings:
- C1 (CRITICAL): force lotusDenoiseBase to same-origin before it reaches
audioWorklet.addModule()/fetch — a crafted call-link param could
otherwise load attacker JS/WASM as a worklet processing the live mic.
Non-same-origin/malformed values fall back to bundled ./denoise/.
- H1 (HIGH): gate audio-inject behind explicit lotusAudioInject=1 (still
acks the action so no transport hang) — it publishes under the local
user's identity, so it must not be silently armed for every call.
- M1 (MED): cap the decoration roster at 512 entries.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Implements RNNoise/Speex noise suppression as a LiveKit audio
TrackProcessor attached to the local mic track, replacing the host's
build-time getUserMedia monkeypatch. Because EC re-attaches the processor
on every (re)publish (LocalTrackPublished), denoise now survives EC's
mid-call reconnect — the root cause of A7 "mic dead after reconnect".
Reuses the worklet/wasm assets already shipped under ./denoise/ (no new EC
dependency); model/gate configured via lotusDenoise/lotusModel/lotusGate
URL params. Additive: no-op unless lotusDenoise=ml.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Review found in-call tiles use MediaView->Avatar, not TileAvatar, so the
decoration never rendered in-call (CRITICAL). Move the overlay into
MediaView, gated on the avatar's own visibility (!(video && videoEnabled))
so it never floats over live video; revert the TileAvatar changes.
Also ref-count the io.lotus.decorations registration (one shared handler,
no double-reply) and stop clearing the map on teardown so a transient
remount doesn't drop decorations (HIGH/MED).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds io.lotus.decorations (toWidget): the host pushes a userId->image-URL
map and EC overlays the profile decoration on each tile avatar
(TileAvatar), keyed by userId, with a useSyncExternalStore-backed store.
Makes A6 first-class in-call instead of absent. URLs are validated
https/blob. Additive: no-op unless the host sends decorations.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Apply the encoding patch to ALL simulcast layers, not just encodings[0]:
screenshare publishes simulcast (VP8), so the full-res layer (the real
bandwidth hog) was left uncapped (CRITICAL).
- Re-apply on TrackUnmuted/restart + a 500ms settle, since LiveKit's
refreshSenderEncodings() overwrites our caps on replaceTrack (device/
source switch, processor toggle) without firing LocalTrackPublished (HIGH).
- Clamp values to sane ranges so a typo can't brick the encoder (MED).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- resume() the AudioContext (host postMessage isn't a gesture) so the clip
isn't silent; warn if it stays suspended (HIGH).
- Close the AudioContext on decode failure (no context leak) (MED).
- Abort in-flight clips on teardown (unmount/vm-change/leave) so audio
doesn't keep blasting to peers (MED).
- Stop the cloned MediaStreamTrack when a room publish fails (MED).
- Validate url is https/blob and fetch with credentials:omit, mode:cors
(MED security).
- Guard against NaN clip duration; fix stale enum doc comment.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds io.lotus.set_quality (toWidget): caps mic audio bitrate and
screenshare bitrate/framerate via RTCRtpSender.setParameters (no
republish). Settings are sticky and re-applied on LocalTrackPublished so
they survive mute/unmute and reconnects. These encoding controls lived in
EC's module scope, unreachable from the host against the prebuilt bundle.
Additive: no-op unless the host sends the action.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds io.lotus.inject_audio (toWidget): mixes a soundboard clip into the
call so other participants hear it. Publishes the clip as a separate
Unknown-source LiveKit track (rendered by MatrixAudioRenderer) rather than
splicing into the mic track, so the denoise pipeline is untouched; the
track is unpublished when the clip ends (with a 30s safety cap). This is
the real call-audio injection that was impossible against the prebuilt EC
bundle. Additive: no-op unless the host sends the action.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Document that io.lotus.call_state is request/response and the host must
ack it (cinny listenAction replies {}) to avoid 10s-timeout churn (H1).
- Throttle 150ms -> 250ms to reduce widget traffic (M1).
- lotusParam: hash fragment wins over query, matching EC's ParamParser (L1).
- Fix the misleading "opaque" id comment; id is userId:deviceId (L2).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds io.lotus.focus_participant (toWidget): the host can pin a participant
to the spotlight by Matrix user id (or clear with userId:null), via a
manual override injected into CallViewModel.spotlightSpeaker$. Replaces
cinny's fragile DOM .click() tile-selector focus hack. Extracts the action
enum into lotusActions.ts (no circular import) and allow-lists Lotus
toWidget actions in initializeWidget. Additive: no-op unless the host
sends the action.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Opt-in (lotusCallState=1) bridge that emits io.lotus.call_state with each
participant's speaking/audio/video state, so the Lotus host can drive
speaking rings / mute badges / PiP from real events instead of scraping
EC's rendered DOM. Exposes vm.userMedia$ on the public CallViewModel.
Additive: no-op without the flag.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>