Desktop-first VTuber assistant with voice, memory, GUI, XML tools, media generation, and Live2D integration.
🇺🇸 English · 🇧🇷 Português (BR)
Lira AM Amarinth is a VTuber assistant designed to run on the Windows desktop, converse via voice, remember context, control local tools, and function as a living character around your PC.
The project is more than just a chatbot: it combines a voice runtime, a GUI panel, hybrid memory, image/music generation, file analysis, PC control via XML tags, VTube Studio integration, and multiple LLM/TTS providers.
The terminal is the "live" mode: short, speakable, and TTS-economical responses. The GUI is the "full chat" mode: longer responses, attachments, markdown, and inline media.
![]() Full Body |
![]() Corinthians Jersey |
![]() Formal Wear |
![]() Beach Outfit |
![]() Control Center |
![]() Visual Chat |
![]() Tools & Settings |
![]() Personalization & UI |
| Area | Capabilities |
|---|---|
| Voice | STT via Whisper, TTS via F5-TTS / XTTS v2 (Colab), ElevenLabs, Google, Azure, OpenAI and Edge, global stop via F8 |
| Agency | Autonomous ReAct Loops (Think-Action-Observe), proactive reminders and task scheduling |
| Brain | Providers google_cloud, groq, openrouter, openai and cerebras |
| GUI | Control Center in CustomTkinter, visual chat, attachments, markdown, audio cards and media preview |
| Multimodality | Video Understanding (frame extraction), image generation/editing, screen awareness and OCR |
| Memory | Chronological SQLite, Vector RAG and Knowledge Graph |
| Desktop | XML tools to open URLs, open apps, read files, move mouse, type, control volume and processes |
| VTuber | VTube Studio integration, emotions, parameters, lipsync and persisted state |
| Discord & WhatsApp | Full-featured Discord Bot (AFK, Auto-mod, Birthdays, Custom commands, Giveaways, Notepad, Reaction Roles, Server Stats, Setup, Sticky messages, Suggestions, support Tickets, Welcome messages) & WhatsApp bridge with TTS & media downloads |
| MCP Gateway | Local HTTP gateway (:8045) for Tavily, GitHub, Filesystem, Memory and Puppeteer — owner-gated + gem economy for public web search |
| Lira Reflex | Owner-only architecture self-inspection: reads docs/ARCHITECTURE.md, roasts the codebase, proposes refactors via MCP filesystem |
Lira now integrates directly with Discord and WhatsApp to act as a community manager and interactive assistant:
- Discord Bot Modules (Cogs):
- AFK: Manage member away statuses with automatic mentions.
- Auto-moderation: Filter and manage server content automatically.
- Birthdays: Set up and announce member birthdays.
- Custom Commands: Create and run custom text commands dynamically.
- Giveaways: Conduct server giveaways with interactive buttons.
- Notepad: Allow users to save and view notes directly through Lira.
- Reaction Roles: Assign server roles via message reactions.
- Server Stats: Display real-time server statistics (member count, channels).
- Setup: Automated guild setup for channels and permissions.
- Sticky Messages: Keep important messages stuck at the bottom of channels.
- Suggestions: Interactive channel for submitting and voting on server suggestions.
- Tickets: Support system using Discord buttons and private channels.
- Welcome: Custom welcome messages with profile cards for new members.
- WhatsApp Bridge: Connects Lira to WhatsApp using Baileys with TTS response support and media downloads (Instagram, TikTok, YouTube, etc.).
- Gem economy (Discord / WhatsApp):
/daily,/weekly,/gemas,/loja_gemas(Discord) — Tavily web search costs 1 gem for public users; owner and Control Panel are free.
| Service | Port | Role |
|---|---|---|
| Control API | 8042 |
GUI chat WebSocket, platform management |
| WhatsApp API | 8043 |
FastAPI chat endpoint for the Baileys bridge |
| Bridge push | 8044 |
Outbound message push from bridge |
| MCP Gateway | 8045 |
Local MCP tool proxy (Tavily, GitHub, filesystem, memory, puppeteer) |
Start the MCP Gateway from Platforms → MCP Gateway in the Control Center, or:
python apps/mcp_gateway/main.pyLira can call external tools silently via XML tags. The gateway runs locally and enforces owner-only access for sensitive servers.
| MCP server | Who can use |
|---|---|
filesystem |
Owner only (Lira Reflex — read project source) |
github |
Owner only |
memory |
Owner only |
puppeteer |
Owner only |
tavily |
Public with gems (1 gem/search); owner & panel = free |
Lira Reflex v1 (owner-only): Lira inspects her own architecture, reads docs/ARCHITECTURE.md, critiques documented tech debt, and can persist architectural notes with <salvar_memoria>. Public users never receive Reflex instructions in the prompt, and the runtime blocks <salvar_memoria> and auto file reads for non-owners.
<mcp>filesystem/read_text_file
{"path": "docs/ARCHITECTURE.md"}</mcp>
<mcp>tavily/tavily_search
latest news about VTuber AI</mcp>Owner identity is resolved via .env:
DISCORD_OWNER_ID=your_discord_snowflake
WPP_OWNER_JID=5511999999999@s.whatsapp.net
WPP_OWNER_LID=your_lid@lid
MCP_GATEWAY_URL=http://127.0.0.1:8045Full architecture manual: docs/ARCHITECTURE.md
Main file: main.py
The terminal is Lira's actual voice mode. It listens to the microphone, builds context, calls the LLM in streaming mode, filters out silent tags, and speaks the response via TTS.
Terminal rule: responses are short by default to avoid long speech and unnecessary TTS costs.
Main file: src/gui/lira_gui.py
The GUI is the complete control panel. In addition to accepting large texts, attachments, drag-and-drop, and media, it features a Platforms (Plataformas) tab to manage, toggle, and view real-time logs for the Discord Bot, WhatsApp API, and Baileys Bridge with a single click, without touching the terminal.
GUI rule: it can provide larger responses when it makes sense, as it acts as a visual chat.
flowchart LR
User["Amarinth"] --> STT["STT Whisper"]
STT --> Runtime["Terminal Runtime"]
User --> GUI["Lira Control Center"]
GUI --> Chat["Visual Chat + Attachments"]
Runtime --> Prompt["Prompt Builder"]
Chat --> Prompt
Prompt --> Memory["Hybrid Memory"]
Memory --> SQLite["SQLite"]
Memory --> RAG["Vector RAG"]
Memory --> Graph["Knowledge Graph"]
Prompt --> LLM["Provider Selector"]
LLM --> Google["Google Gemini / Vertex"]
LLM --> Groq["Groq"]
LLM --> OpenRouter["OpenRouter"]
LLM --> OpenAI["OpenAI"]
LLM --> Cerebras["Cerebras"]
LLM --> Divider["SentenceDivider"]
Divider --> TTS["TTS Selector"]
TTS --> F5Colab["F5-TTS / XTTS (Colab)"]
TTS --> Eleven["ElevenLabs"]
TTS --> GoogleTTS["Google TTS"]
TTS --> Azure["Azure"]
TTS --> Edge["Edge"]
Divider --> XML["Silent XML Tags"]
XML --> Media["Image / Music"]
XML --> PC["PC Control"]
XML --> VTS["VTube Studio"]
Lira can say one thing to the user and execute another in the background, without leaking the command to the TTS.
<salvar_memoria>Amarinth prefers short answers in the terminal.</salvar_memoria>
<gerar_imagem>anime portrait of Lira with golden hair and white flowers</gerar_imagem>
<acao_pc>{"action":"open_url","url":"https://github.com"}</acao_pc>
<analisar_video>https://youtube.com/watch?v=ID</analisar_video>
<agendar_aviso>{"time": "10m", "message": "Take the cake out of the oven"}</agendar_aviso>
<mcp>tavily/tavily_search
open source VTuber assistants</mcp>Important rules:
- XML tags only apply to the current request;
- Past actions should not be repeated after a long pause;
run_commandand dangerous operations go through guardrails;- Text inside the tags must not be spoken by the TTS;
<salvar_memoria>, Lira Reflex and MCP filesystem/github/memory/puppeteer are owner-only — blocked at gateway and runtime for everyone else.
Current fallback chain:
f5_colab -> edge -> elevenlabs -> google -> azure -> openai
The F8 hotkey triggers a global stop to attempt to halt:
- Terminal speech;
- GUI chat speech;
- Audio previews;
- Music or playback started by Lira.
Lira's memory uses three layers:
| Layer | Function |
|---|---|
| SQLite | Recent chronological history |
| RAG | Semantic search for old conversations |
| Knowledge Graph | Permanent facts and relationships |
The terminal also receives temporal context: if too much time has passed since the last speech, Lira treats the new message as a new context and avoids continuing old tasks automatically.
Lira's development is organized into phases focusing on fluidity, immersion, and computer vision:
- Autonomous Tool Loops: Sequential tool use in a single turn.
- Lira Scheduler: Proactive reminders and notifications.
- Self-Memory Maintenance: Automated cleanup of old context.
- Video Understanding: Frame extraction and visual commentary.
- Voice-to-Voice (WA): Native audio processing for WhatsApp.
- Multimodal OCR: Deep reading of complex documents and screens.
- Emotional RVC Mapping: Dynamic pitch adjustment based on mood.
- Extended VTS Bridge: Real-time parameter sync via WebSockets.
- Episodic Memory: Daily summaries of user interactions.
Lira AM Amarinth e uma assistente VTuber feita para rodar no desktop do Windows, conversar por voz, lembrar contexto, controlar ferramentas locais e funcionar como uma personagem viva em volta do seu PC.
O projeto nao e so um chatbot: ele combina runtime de voz, painel GUI, memoria hibrida, geracao de imagem/musica, analise de arquivos, controle do PC por tags XML, VTube Studio e multiplos providers de LLM/TTS.
O terminal e o modo "ao vivo": respostas curtas, falaveis e economicas para TTS. A GUI e o modo "chat completo": respostas maiores, anexos, markdown, arquivos e midia inline.
![]() Corpo Inteiro |
![]() Camisa do Timão |
![]() Roupa Social |
![]() Visual Praia |
![]() Control Center |
![]() Chat Visual |
![]() Ferramentas & Configs |
![]() Personalização e Cores |
| Area | Recursos |
|---|---|
| Voz | STT por Whisper, TTS via F5-TTS / XTTS v2 (Colab), ElevenLabs, Google, Azure, OpenAI e Edge, stop global por F8 |
| Agência | Loops ReAct Autônomos (Pensa-Age-Observa), lembretes proativos e agendamentos |
| Cerebro | Providers google_cloud, groq, openrouter, openai e cerebras |
| GUI | Control Center em CustomTkinter, chat visual, anexos, markdown, cards de audio e preview de midia |
| Multimodalidade | Entendimento de Vídeo (extração de frames), geração de imagem, visão de tela e OCR |
| Memoria | SQLite cronologico, RAG vetorial e grafo de conhecimento |
| Desktop | Tool XML para abrir URL, abrir app, ler arquivo, mover mouse, digitar, volume e processos |
| VTuber | Integracao com VTube Studio, emocoes, parametros, lipsync e estado persistido |
| Discord & WhatsApp | Bot de Discord completo (AFK, Auto-moderação, Aniversários, Comandos customizados, Sorteios, Bloco de notas, Cargos por reação, Status do servidor, Setup automatizado, Mensagens fixas, Sugestões, Tickets de suporte, Boas-vindas) & Ponte para WhatsApp com suporte a TTS e downloads de mídia |
| MCP Gateway | Gateway HTTP local (:8045) para Tavily, GitHub, Filesystem, Memory e Puppeteer — restrito ao criador + economia de gemas na busca web |
| Lira Reflex | Autoconsulta de arquitetura exclusiva do criador: lê docs/ARCHITECTURE.md, critica o código e propõe refactors via MCP filesystem |
A Lira agora se conecta diretamente ao Discord e ao WhatsApp para atuar como moderadora, gerente de comunidade e assistente interativa:
- Módulos do Bot do Discord (Cogs):
- AFK: Gerenciamento de status ausente de membros com menções automáticas.
- Auto-moderação: Filtro automático de conteúdo e moderação de chat.
- Aniversários: Registro e anúncio automático de aniversários de membros.
- Comandos Customizados: Criação e execução de comandos de texto personalizados de forma dinâmica.
- Sorteios (Giveaways): Realização de sorteios interativos com botões.
- Bloco de Notas (Notepad): Permite que usuários salvem anotações rápidas pela Lira.
- Cargos por Reação (Reaction Roles): Atribuição de cargos no servidor através de reações a mensagens.
- Status do Servidor: Exibição de estatísticas do servidor em tempo real.
- Setup: Configuração guiada e automatizada de canais e permissões.
- Mensagens Fixas (Sticky Messages): Mantém mensagens importantes sempre no fim do chat.
- Sugestões: Canal interativo para envio e votação de sugestões dos membros.
- Tickets: Sistema de suporte/atendimento privado via botões do Discord.
- Boas-vidas: Mensagens personalizadas de entrada com geração de cartões de perfil para novos membros.
- Ponte do WhatsApp: Conecta a Lira ao WhatsApp usando a biblioteca Baileys, com respostas em áudio (TTS) e downloads de mídia (Instagram, TikTok, YouTube, etc.).
- Economia de gemas (Discord / WhatsApp):
/daily,/weekly,/gemas,/loja_gemas(Discord) — busca Tavily custa 1 gema para usuários públicos; criador e painel são grátis.
| Serviço | Porta | Função |
|---|---|---|
| Control API | 8042 |
Chat WebSocket da GUI, gerenciamento de plataformas |
| WhatsApp API | 8043 |
Endpoint FastAPI de chat para a ponte Baileys |
| Bridge push | 8044 |
Push de mensagens outbound da ponte |
| MCP Gateway | 8045 |
Proxy local de ferramentas MCP (Tavily, GitHub, filesystem, memory, puppeteer) |
Ligue o MCP Gateway em Plataformas → MCP Gateway no Control Center, ou:
python apps/mcp_gateway/main.pyA Lira chama ferramentas externas em silêncio via tags XML. O gateway roda localmente e aplica controle de acesso por criador nos servidores sensíveis.
| Servidor MCP | Quem pode usar |
|---|---|
filesystem |
Só o criador (Lira Reflex — leitura do código-fonte) |
github |
Só o criador |
memory |
Só o criador |
puppeteer |
Só o criador |
tavily |
Público com gemas (1 gema/busca); criador e painel = grátis |
Lira Reflex v1 (só criador): a Lira inspeciona a própria arquitetura, lê docs/ARCHITECTURE.md, debocha das gambiarras documentadas e pode persistir reflexões com <salvar_memoria>. Usuários públicos não recebem instruções de Reflex no prompt, e o runtime bloqueia <salvar_memoria> e leitura automática de arquivos para quem não é dono.
<mcp>filesystem/read_text_file
{"path": "docs/ARCHITECTURE.md"}</mcp>
<mcp>tavily/tavily_search
notícias sobre VTuber IA</mcp>Identidade do criador via .env:
DISCORD_OWNER_ID=seu_id_discord
WPP_OWNER_JID=5511999999999@s.whatsapp.net
WPP_OWNER_LID=seu_lid@lid
MCP_GATEWAY_URL=http://127.0.0.1:8045Manual completo: docs/ARCHITECTURE.md
Arquivo principal: main.py
O terminal e o modo de voz real da Lira. Ele escuta o microfone, monta contexto, chama o LLM em streaming, filtra tags silenciosas e fala a resposta por TTS.
Regra do terminal: respostas curtas por padrao para evitar fala longa e gasto desnecessario de TTS.
Arquivo principal: src/gui/lira_gui.py
A GUI é o painel de controle completo. Além de aceitar textos grandes, anexos, drag-and-drop, Ctrl+V, arquivos de código, imagens, áudio e vídeo, ela conta com a aba Plataformas para gerenciar, ligar, desligar e acompanhar os logs em tempo real do Bot do Discord, da API do WhatsApp e da Ponte Baileys com um clique, sem precisar do terminal.
Regra da GUI: pode responder grande quando fizer sentido, porque funciona como chat visual.
flowchart LR
User["Amarinth"] --> STT["STT Whisper"]
STT --> Runtime["Terminal Runtime"]
User --> GUI["Lira Control Center"]
GUI --> Chat["Chat Visual + Anexos"]
Runtime --> Prompt["Prompt Builder"]
Chat --> Prompt
Prompt --> Memory["Memoria Hibrida"]
Memory --> SQLite["SQLite"]
Memory --> RAG["RAG Vetorial"]
Memory --> Graph["Knowledge Graph"]
Prompt --> LLM["Provider Selector"]
LLM --> Google["Google Gemini / Vertex"]
LLM --> Groq["Groq"]
LLM --> OpenRouter["OpenRouter"]
LLM --> OpenAI["OpenAI"]
LLM --> Cerebras["Cerebras"]
LLM --> Divider["SentenceDivider"]
Divider --> TTS["TTS Selector"]
TTS --> F5Colab["F5-TTS / XTTS (Colab)"]
TTS --> Eleven["ElevenLabs"]
TTS --> GoogleTTS["Google TTS"]
TTS --> Azure["Azure"]
TTS --> Edge["Edge"]
Divider --> XML["Tags XML Silenciosas"]
XML --> Media["Imagem / Musica"]
XML --> PC["Controle do PC"]
XML --> VTS["VTube Studio"]
A Lira pode falar uma coisa para o usuario e executar outra por baixo, sem vazar comando no TTS.
<salvar_memoria>O Amarinth prefere respostas curtas no terminal.</salvar_memoria>
<gerar_imagem>anime portrait of Lira with golden hair and white flowers</gerar_imagem>
<acao_pc>{"action":"open_url","url":"https://github.com"}</acao_pc>
<analisar_video>https://youtube.com/watch?v=ID</analisar_video>
<agendar_aviso>{"tempo": "5m", "mensagem": "Tirar o bolo do forno"}</agendar_aviso>
<mcp>tavily/tavily_search
assistentes VTuber open source</mcp>Regras importantes:
- tags XML so valem para o pedido atual;
- acoes antigas nao devem ser repetidas depois de pausa longa;
run_commande operacoes perigosas passam por guardrails;- o texto dentro das tags nao deve ser falado pelo TTS;
<salvar_memoria>, Lira Reflex e MCP filesystem/github/memory/puppeteer sao exclusivos do criador — bloqueados no gateway e no runtime para todos os outros.
Fallback atual:
f5_colab -> edge -> elevenlabs -> google -> azure -> openai
ElevenLabs suporta configuracao manual pela GUI:
voice_idmodel_idratestabilitysimilarity_booststylespeaker_boost
O hotkey F8 usa stop global para tentar parar:
- fala do terminal;
- fala do chat da GUI;
- preview de audio;
- musica ou playback iniciado pela Lira.
A memoria da Lira usa tres camadas:
| Camada | Funcao |
|---|---|
| SQLite | historico cronologico recente |
| RAG | busca semantica por conversas antigas |
| Knowledge Graph | fatos permanentes e relacionamentos |
O terminal tambem recebe contexto temporal: se passou muito tempo desde a ultima fala, a Lira trata a nova mensagem como novo contexto e evita continuar tarefas antigas automaticamente.
git clone https://github.com/AmarinthIA/AmarinthLira-VTuber-OSS.git
cd AmarinthLira-VTuber-OSSpython -m venv .venv
.venv\Scripts\activatepip install -r requirements.txtDependencias opcionais:
pip install pyvts pypdf PyPDF2 youtube-transcript-apiGuia completo: docs/INSTALL.md
Copie:
copy .env.example .envExemplo:
GROQ_API_KEY=sua_chave
GEMINI_API_KEY=sua_chave
OPENROUTER_API_KEY=sua_chave
OPENAI_API_KEY=sua_chave
CEREBRAS_API_KEY=sua_chave
ELEVENLABS_API_KEY=sua_chave
TAVILY_API_KEY=sua_chave
GITHUB_TOKEN=seu_pat_github
MCP_GATEWAY_URL=http://127.0.0.1:8045
DISCORD_OWNER_ID=seu_id_discord
WPP_OWNER_JID=5511999999999@s.whatsapp.net
WPP_OWNER_LID=seu_lid@lid
GOOGLE_CLOUD_PROJECT=seu_projeto
GOOGLE_CLOUD_LOCATION=globalNunca commite seu .env.
Para o setup minimo, apenas GROQ_API_KEY e obrigatoria. O TTS publico default e edge, sem chave.
python apps/vtuber/main.py(Compat: python main.py — shim para o mesmo runtime.)
python -m src.gui.lira_guicscript //nologo run_lira_gui_hidden.vbspython -m src.modules.discord.bot- Inicie a API dedicada ao WhatsApp:
python apps/whatsapp_api/main.py- Inicie a ponte Node.js:
cd whatsapp_bridge
node index.jsArquivo: src/config/config.json
No clone publico, os defaults ficam em src/config/config.example.json. O arquivo src/config/config.json e local e ignorado pelo Git.
Guias detalhados: docs/CONFIG.md, docs/PROVIDERS.md, docs/TROUBLESHOOTING.md
Blocos importantes:
{
"LLM_PROVIDER": "groq",
"TTS_PROVIDER": "edge",
"CHAT": {
"LLM_PROVIDER": "groq",
"response_mode": "adaptive",
"auto_route_media": true
},
"GUI": {
"stop_hotkey_enabled": true,
"stop_hotkey": "F8"
}
}| Provider | Uso |
|---|---|
google_cloud |
Gemini API, Vertex AI, visao e midia |
groq |
respostas rapidas e modelos open |
openrouter |
acesso a varios modelos por gateway |
openai |
modelos OpenAI diretos |
cerebras |
inferencia rapida quando configurado |
O provider Google trabalha em modo:
gemini_apivertex_aiauto
A integracao fica em src/modules/vts_controller.py.
Ela inclui:
- autenticacao por token;
- heartbeat;
- reconexao;
- leitura de hotkeys;
- expressoes;
- parametros;
- tentativa de lipsync;
- estado salvo em
data/vts_state.json.
Pasta padrao:
%USERPROFILE%\Desktop\lira_inbox
Estrutura:
lira_inbox/
imagem/
pdf/
docs/
code/
audio/
video/
A GUI tambem aceita arrastar arquivos direto no chat, entao a inbox e util, mas nao obrigatoria.
python apps/vtuber/main.py
python apps/control_api/main.py
python apps/mcp_gateway/main.py
python apps/whatsapp_api/main.py
python -m src.modules.discord.bot
python -m compileall apps main.py src
python -m pytest -qO desenvolvimento da Lira está organizado em fases para focar em fluidez, imersão e visão computacional:
- Loops de Ferramentas Autônomos: Uso sequencial de ferramentas em um único turno.
- Lira Scheduler: Lembretes e notificações proativas.
- Manutenção de Auto-Memória: Limpeza automatizada de contextos antigos.
- Entendimento de Vídeo: Extração de frames e comentários visuais.
- Voz-para-Voz (WhatsApp): Processamento de áudio nativo para o bridge.
- OCR Multimodal: Leitura profunda de documentos e telas complexas.
- Mapeamento Emocional RVC: Ajuste dinâmico de tom de voz baseado no humor.
- Bridge VTS Estendido: Sincronia de parâmetros via WebSockets.
- Memória Episódica: Resumos diários das interações com o usuário.
O projeto ja e funcional, mas ainda esta em evolucao rapida.
Pontos fortes:
- desktop-first;
- GUI e terminal separados por contrato;
- memoria hibrida;
- tags XML silenciosas;
- providers multiplos;
- stop global;
- foco real em VTuber e assistente residente.
Limites atuais:
- foco principal em Windows;
- alguns providers exigem chave paga;
- Live2D depende de modelo e parametros corretos;
- algumas automacoes desktop ainda precisam de mais guardrails.
Contribuicoes devem preservar a separacao entre:
- terminal de voz;
- chat visual da GUI;
- memoria;
- providers;
- TTS;
- tags XML;
- controle do PC.
Se mudar comportamento de prompt, canal ou provider, atualize a documentacao junto.
Consulte a licenca do repositorio e os termos dos providers externos usados.







