Skip to content

中文版

Generated reference snapshot

Original Cubic page · captured 2026-09-23 · source commit. This AI-generated page has not been verified against the current code. Use the Zhin documentation for current behavior and see known corrections.

Known correction

Some source citation line anchors in the Cubic original point to unrelated lines at the pinned commit. Check the linked file rather than relying on its line range. See Speech.

Relevant source files

The following files were used as context for generating this wiki page:

Speech Pipeline (STT & TTS)

The Speech Pipeline in Zhin.js provides optional Speech-to-Text (STT) and Text-to-Speech (TTS) capabilities. It enables multi-channel bots to process inbound voice data and generate outbound audio responses. The system operates as a modular "tier" that you can install only when the product requirements necessitate rich media interaction.

Sources: README.md:92-99, README.zh-CN.md:123

Architecture and Data Flow

The speech module integrates directly into the Zhin.js message pipeline. When active, it intercepts inbound audio segments for transcription and converts outbound text segments into playable audio formats.

The diagram shows how speech segments flow between external adapters and the internal message pipeline via the speech toolkit. Sources: README.md:83-91, packages/toolkit/speech/package.json:20-22

Tier-Based Capabilities

Speech services are categorized under the "Connect" surface of the capability map. If the @zhin.js/speech package is missing, the system issues a warning and degrades to standard text processing.

Sources: README.md:104-106, README.zh-CN.md:123

Components and Providers

The speech pipeline supports multiple industry-standard engines. These are managed via the @zhin.js/speech package, which lists core speech keywords such as Whisper and Edge-TTS.

FeatureDescriptionSupported Engines
STTInbound Speech-to-Text transcription.Whisper, Custom providers
TTSOutbound Text-to-Speech synthesis.Edge-TTS, OpenAI, Azure, Custom
SegmentsAudio segment handling in the message pipeline.segment.tts

Sources: packages/toolkit/speech/package.json:20-22, README.zh-CN.md:123

Installation and Integration

The speech module is an advanced capability that adds approximately a few megabytes to the production size. It is not included in the default <10MB IM core installation.

Dependency Management

To enable speech features, you must add the speech toolkit to your project:

bash
pnpm add @zhin.js/speech

The package requires @zhin.js/core as a peer dependency and is written in TypeScript using ESM modules.

Sources: README.md:126, packages/toolkit/speech/package.json:1-12, packages/toolkit/speech/package.json:28-34

Integration Sequence

The following sequence illustrates a typical speech interaction turn:

Sources: README.md:83-91, README.zh-CN.md:123

Configuration Summary

While primary configuration for individual providers is handled through standard Zhin.js configuration projection, the speech module itself is defined as a toolkit package.

Configuration ElementRoleReference File
@zhin.js/speechEntry point for speech capabilities.packages/toolkit/speech/package.json
segment.ttsThe primary UI component for triggering synthesis.README.md
Connect SurfaceGoverns adapters and speech integration rules.README.md

Sources: README.md:104-106, packages/toolkit/speech/package.json:1-5

Conclusion

The Zhin.js Speech Pipeline provides a scalable solution for audio-based bot interactions. By separating STT and TTS into the @zhin.js/speech toolkit, the framework maintains a lightweight core while offering deep integration with professional audio synthesis and transcription engines like Whisper and Edge-TTS. If the module is not installed, the framework ensures stability through graceful degradation to text.

Sources: README.md:92-99, README.zh-CN.md:123