livekit-plugins-speechmatics 1.8.3


pip install livekit-plugins-speechmatics

  Latest version

Released: Sep 23, 2026


Meta
Author: LiveKit
Requires Python: >=3.10.0

Classifiers

Intended Audience
  • Developers

License
  • OSI Approved :: Apache Software License

Programming Language
  • Python :: 3
  • Python :: 3 :: Only
  • Python :: 3.10

Topic
  • Multimedia :: Sound/Audio
  • Multimedia :: Video
  • Scientific/Engineering :: Artificial Intelligence

Speechmatics STT plugin for LiveKit Agents

Support for Speechmatics STT.

See https://docs.livekit.io/agents/integrations/stt/speechmatics/ for more information.

Installation

pip install livekit-plugins-speechmatics

Model

model selects the transcription model and defaults to linden-1, currently the only Agent STT model. Pass it as a string or as speechmatics.Model.LINDEN_1.

operating_point is a deprecated alias for model and warns when used. The RT operating points enhanced and standard are not Agent STT models and are rejected by the service.

Turn detection modes

The turn_detection_mode parameter controls how end-of-turn (endpointing) is detected:

  • VAD (default) — Speechmatics runs its own VAD and closes turns itself (service-side endpointing). No vad is required. Pair it with turn_detection="stt" on the AgentSession, otherwise the session's own turn detector decides and Speechmatics' end-of-turn is ignored.
  • EXTERNAL — Speechmatics does not endpoint on its own. Turns close when the caller calls finalize(). In practice you pass a vad to the plugin and its end-of-speech drives finalize(); LiveKit does not call finalize() for you, and no VAD is auto-loaded. Without a vad (and without calling finalize() yourself) turns never close, so nothing is finalized. The session's own turn detector decides when the user's turn ends; finalize() only makes Speechmatics flush what it has as a final segment.

The earlier FIXED, ADAPTIVE and SMART_TURN modes each selected one of the old engine's service-side endpointing strategies. Agent STT exposes a single one, so all three are deprecated and resolve to VAD with a warning. FIXED additionally loses its end_of_utterance_silence_trigger timing, which Agent STT does not support.

Usage — service-side endpointing (VAD, default)

Let Speechmatics detect turns and tell the session to act on them:

from livekit.agents import AgentSession
from livekit.plugins import speechmatics

agent = AgentSession(
    stt=speechmatics.STT(),
    turn_detection="stt",
    ...
)

Usage — caller-driven endpointing (EXTERNAL)

Pass a vad to the plugin; its end-of-speech drives finalize(). AgentSession loads its own VAD when none is given, so pass the same instance to both and a single VAD serves the session and the plugin:

from livekit.agents import AgentSession, inference
from livekit.plugins import speechmatics

vad = inference.VAD()

agent = AgentSession(
    stt=speechmatics.STT(
        turn_detection_mode=speechmatics.TurnDetectionMode.EXTERNAL,
        # The VAD passed here drives finalize() on end-of-speech.
        vad=vad,
        speaker_format="[Speaker {speaker_id}] {text}",
    ),
    vad=vad,
    ...
)

Interim transcripts

The service sends partial segments by default, emitted as interim transcripts. Set include_partials=False to receive final segments only:

stt = speechmatics.STT(include_partials=False)

Diarization

Speechmatics attributes each transcript segment to a speaker. Diarization is enabled by default (enable_diarization=True); the segment is the unit of attribution, so each result carries a single speaker_id and there is no per-word speaker data. To fold the speaker label into the transcript text, set speaker_format using the {speaker_id} and {text} placeholders:

  • speaker_format="<{speaker_id}>{text}</{speaker_id}>" -> <S1>Hello</S1>
  • speaker_format="[Speaker {speaker_id}] {text}" -> [Speaker S1] Hello

Segments the service did not attribute — including every segment when diarization is off — are labelled UU.

Adjust your system instructions to inform the LLM of this format so it can attribute speakers.

from livekit.agents import AgentSession
from livekit.plugins import speechmatics

agent = AgentSession(
    stt=speechmatics.STT(
        enable_diarization=True,
        max_speakers=4,
        speaker_format="[Speaker {speaker_id}] {text}",
        additional_vocab=[
            speechmatics.AdditionalVocabEntry(
                content="LiveKit",
                sounds_like=["live kit"],
            ),
        ],
    ),
    ...
)

Speaker identification

Speaker labels are per session by default, so S1 in one session is unrelated to S1 in the next. To carry labels across sessions, read the identifiers out of a live session with await stt.get_speaker_ids() (call it once each speaker has said a few words), then pass them back as known_speakers on a later session:

speakers = await stt.get_speaker_ids()

stt = speechmatics.STT(known_speakers=speakers)

Each entry maps a human-readable label to the identifiers the engine recognizes it by. With more than one open stream the call returns one list per stream instead.

Pre-requisites

You'll need to specify a Speechmatics API Key. It can be set as environment variable SPEECHMATICS_API_KEY or in a .env.local file.

The plugin connects to wss://eu2.rt.speechmatics.com/v2/agent by default. To use another region or a self-hosted endpoint, set base_url or the SPEECHMATICS_RT_URL environment variable.

1.8.3 Sep 23, 2026
1.8.2 Sep 15, 2026
1.8.1 Sep 10, 2026
1.8.0 Sep 05, 2026
1.7.1 Aug 27, 2026
1.7.0 Aug 20, 2026
1.6.11rc1 Aug 13, 2026
1.6.10 Aug 13, 2026
1.6.9 Aug 07, 2026
1.6.8 Aug 03, 2026
1.6.7 Jul 25, 2026
1.6.6 Jul 18, 2026
1.6.5 Jul 09, 2026
1.6.4 Jun 24, 2026
1.6.3 Jun 22, 2026
1.6.2 Jun 19, 2026
1.6.1 Jun 17, 2026
1.6.0 Jun 11, 2026
1.6.0rc2 May 29, 2026
1.6.0rc1 May 27, 2026
1.5.19rc1 Jun 08, 2026
1.5.18 Jun 05, 2026
1.5.17 Jun 03, 2026
1.5.16 Jun 01, 2026
1.5.15 May 29, 2026
1.5.14 May 27, 2026
1.5.13 May 25, 2026
1.5.12 May 21, 2026
1.5.11 May 19, 2026
1.5.10 May 18, 2026
1.5.9 May 13, 2026
1.5.8 May 05, 2026
1.5.7 Apr 30, 2026
1.5.6 Apr 22, 2026
1.5.5 Apr 20, 2026
1.5.4 Apr 16, 2026
1.5.3 Apr 15, 2026
1.5.2 Apr 08, 2026
1.5.1 Mar 23, 2026
1.5.0 Mar 19, 2026
1.5.0rc2 Mar 06, 2026
1.5.0rc1 Feb 13, 2026
1.4.6 Mar 16, 2026
1.4.5 Mar 11, 2026
1.4.4 Mar 03, 2026
1.4.3 Feb 23, 2026
1.4.2 Feb 17, 2026
1.4.1 Feb 06, 2026
1.4.0rc2 Jan 23, 2026
1.4.0rc1 Dec 23, 2025
1.3.12 Jan 21, 2026
1.3.11 Jan 14, 2026
1.3.10 Dec 23, 2025
1.3.9 Dec 19, 2025
1.3.8 Dec 17, 2025
1.3.7 Dec 16, 2025
1.3.6 Dec 03, 2025
1.3.5 Nov 25, 2025
1.3.4 Nov 24, 2025
1.3.3 Nov 19, 2025
1.3.2 Nov 17, 2025
1.3.1 Nov 17, 2025
1.3.0rc2 Nov 15, 2025
1.3.0rc1 Nov 06, 2025
1.2.18 Nov 05, 2025
1.2.17 Oct 29, 2025
1.2.16 Oct 27, 2025
1.2.15 Oct 15, 2025
1.2.14 Oct 01, 2025
1.2.13 Oct 01, 2025
1.2.12 Sep 29, 2025
1.2.11 Sep 18, 2025
1.2.9 Sep 15, 2025
1.2.8 Sep 02, 2025
1.2.7 Aug 28, 2025
1.2.6 Aug 18, 2025
1.2.5 Aug 10, 2025
1.2.4 Aug 07, 2025
1.2.3 Aug 04, 2025
1.2.2 Jul 24, 2025
1.2.1 Jul 17, 2025
1.2.0 Jul 17, 2025
1.1.7 Jul 15, 2025
1.1.6 Jul 10, 2025
1.1.5 Jun 30, 2025
1.1.4 Jun 25, 2025
1.1.3 Jun 21, 2025
1.1.2 Jun 20, 2025
1.1.1 Jun 10, 2025
1.1.0 Jun 10, 2025
1.0.23 May 29, 2025
1.0.22 May 17, 2025
1.0.21 May 15, 2025
1.0.20 May 08, 2025
1.0.19 May 03, 2025
1.0.18 May 01, 2025
1.0.17 Apr 24, 2025
1.0.16 Apr 22, 2025
1.0.15 Apr 22, 2025
1.0.14 Apr 22, 2025
1.0.13 Apr 15, 2025
1.0.12 Apr 15, 2025
1.0.11 Apr 10, 2025
1.0.0rc9 Apr 07, 2025
1.0.0rc8 Apr 07, 2025
1.0.0rc7 Apr 07, 2025
1.0.0rc6 Apr 03, 2025
1.0.0rc5 Apr 03, 2025
1.0.0rc4 Mar 29, 2025
1.0.0rc3 Mar 27, 2025
1.0.0rc2 Mar 27, 2025
1.0.0rc1 Mar 26, 2025
1.0.0.dev5 Mar 19, 2025
1.0.0.dev4 Mar 19, 2025
0.0.3 May 19, 2025
0.0.2 Mar 06, 2025
Extras: None
Dependencies:
livekit-agents (>=1.8.3)
speechmatics-agent-stt (~=0.1.0)
speechmatics-voice (>=0.2.8)