CTC segmentation to align utterances within large audio files (ESPnet fork).
Project Links
Meta
Author: Ludwig Kuerzinger <ludwig.kuerzinger@tum.de>, Dominik Winkelbauer <dominik.winkelbauer@tum.de>
Maintainer: ESPnet Developers
Requires Python: >=3.9
Classifiers
CTC segmentation
This is an unofficial republication of
lumaku/ctc-segmentation maintained
by the ESPnet project. It is not affiliated with, nor endorsed by, the original
authors. It exists only because upstream's released 1.7.4 builds against the
NumPy 1.x ABI and so fails to import under NumPy 2 - a problem upstream has
already fixed on master but has not released. The import name is unchanged.
See FORK_NOTICE.md.
CTC segmentation is used to align utterances within audio files. It can be combined with CTC-based ASR models. This package includes the core functions.