From 05219a38e375e12bcab65ccfdba6a872399034af Mon Sep 17 00:00:00 2001 From: rfshubert Date: Sun, 22 Feb 2026 08:53:53 -0300 Subject: [PATCH] docs: add Voice Transcription (STT) section to README Document the stt_model configuration field and how to use any OpenAI-compatible Whisper endpoint for voice transcription. Update the Providers note to reflect the generalized STT support. --- README.md | 54 ++++++++++++++++++++++++++++++++++++++++++++++++++++-- 1 file changed, 52 insertions(+), 2 deletions(-) diff --git a/README.md b/README.md index de6fd87ea..397c314c2 100644 --- a/README.md +++ b/README.md @@ -767,7 +767,7 @@ The subagent has access to tools (message, web_search, etc.) and can communicate ### Providers > [!NOTE] -> Groq provides free voice transcription via Whisper. If configured, Telegram voice messages will be automatically transcribed. +> PicoClaw supports voice transcription (STT) on Telegram, Discord, Slack, and OneBot channels. You can use **any OpenAI-compatible Whisper endpoint** (OpenAI, Groq, local Whisper servers, etc.) by configuring `stt_model` in `agents.defaults`. See [Voice Transcription (STT)](#voice-transcription-stt) for details. | Provider | Purpose | Get API Key | | -------------------------- | --------------------------------------- | -------------------------------------------------------------------- | @@ -778,7 +778,7 @@ The subagent has access to tools (message, web_search, etc.) and can communicate | `openai(To be tested)` | LLM (GPT direct) | [platform.openai.com](https://platform.openai.com) | | `deepseek(To be tested)` | LLM (DeepSeek direct) | [platform.deepseek.com](https://platform.deepseek.com) | | `qwen` | LLM (Qwen direct) | [dashscope.console.aliyun.com](https://dashscope.console.aliyun.com) | -| `groq` | LLM + **Voice transcription** (Whisper) | [console.groq.com](https://console.groq.com) | +| `groq` | LLM + STT (Whisper) | [console.groq.com](https://console.groq.com) | | `cerebras` | LLM (Cerebras direct) | [cerebras.ai](https://cerebras.ai) | ### Model Configuration (model_list) @@ -1093,6 +1093,56 @@ picoclaw agent -m "Hello" +### Voice Transcription (STT) + +PicoClaw can automatically transcribe voice messages on Telegram, Discord, Slack, and OneBot channels. Instead of being limited to a single hardcoded provider, you can use **any OpenAI-compatible Whisper endpoint** by setting the `stt_model` field in `agents.defaults`. + +#### Configuration + +Add a Whisper-compatible model to your `model_list` and point `stt_model` to it: + +```json +{ + "agents": { + "defaults": { + "model": "my-llm", + "stt_model": "whisper" + } + }, + "model_list": [ + { + "model_name": "whisper", + "model": "openai/whisper-1", + "api_key": "sk-your-openai-key", + "api_base": "https://api.openai.com/v1" + } + ] +} +``` + +This works with any provider that exposes the `/audio/transcriptions` endpoint (OpenAI, Groq, local Whisper servers, etc.): + +| Provider | `model` value | Notes | +| --- | --- | --- | +| **OpenAI** | `openai/whisper-1` | High accuracy | +| **Groq** | `groq/whisper-large-v3` | Fast inference, free tier | +| **Local** | `openai/whisper-1` | Set `api_base` to your local server | + +#### Backward Compatibility + +If `stt_model` is not set, PicoClaw falls back to legacy Groq detection: + +1. `providers.groq.api_key` — uses Groq Whisper automatically +2. Any `groq/` entry in `model_list` — uses that entry's API key for Groq Whisper + +No existing configuration will break. + +#### Environment Variable + +```bash +export PICOCLAW_AGENTS_DEFAULTS_STT_MODEL=whisper +``` + ## CLI Reference | Command | Description |