diff --git a/mintlify-docs/ai-tools/build-local-rgb-agent.mdx b/mintlify-docs/ai-tools/build-local-rgb-agent.mdx index 619f017..baca55c 100644 --- a/mintlify-docs/ai-tools/build-local-rgb-agent.mdx +++ b/mintlify-docs/ai-tools/build-local-rgb-agent.mdx @@ -108,11 +108,11 @@ npm init -y && npm pkg set type=module npm install @kaleidorg/mind @qvac/sdk @modelcontextprotocol/sdk ``` -KaleidoMind works with `@qvac/sdk` 0.13 and later; installing the latest release is recommended. Save this as `agent.mjs`: +KaleidoMind works with `@qvac/sdk` 0.13.1 and later and is tested with 0.21. Save this as `agent.mjs`: ```js agent.mjs import { createInterface } from 'node:readline/promises'; -import { completion, cancel, loadModel, QWEN3_600M_INST_Q4 } from '@qvac/sdk'; +import { completion, cancel, loadModel, QWEN3_5_2B_MULTIMODAL_Q4_K_M } from '@qvac/sdk'; import { Funnel, ToolRegistry, confirmReadback } from '@kaleidorg/mind'; import { McpToolSource } from '@kaleidorg/mind/mcp'; import { createQvacProvider } from '@kaleidorg/mind/qvac'; @@ -136,11 +136,15 @@ await kaleido.connect(); // 2. A local model through QVAC (downloaded on first run) const modelId = await loadModel({ - modelSrc: QWEN3_600M_INST_Q4, + modelSrc: QWEN3_5_2B_MULTIMODAL_Q4_K_M, modelType: 'llm', modelConfig: { ctx_size: 8192, tools: true }, }); -const provider = createQvacProvider({ completion, cancel, getModelId: () => modelId }); +const provider = createQvacProvider({ + completion, cancel, getModelId: () => modelId, + defaultMaxTokens: 1536, // Qwen3.5 reasons before answering; leave room for both + maxThinkingTokens: 512, +}); // 3. The tiered funnel, with a confirmation prompt before any spend const funnel = new Funnel({ provider, tools: new ToolRegistry([kaleido]) }); @@ -168,7 +172,7 @@ node agent.mjs The first start downloads the model. Every fund-moving tool pauses on the `onConfirm` callback, so nothing leaves the node until you type `y`. - This is a minimal host. The [`@kaleidorg/mind` README](https://www.npmjs.com/package/@kaleidorg/mind) and the `examples/node-minimal` and `examples/rgb-agent` folders in the [kaleido-mind repository](https://github.com/kaleidoswap/kaleido-mind) go further, with skills, recipes, and a larger model. A 0.6B model handles the fast path and the recipes well; for open-ended requests a larger model such as Qwen3 1.7B or 4B does noticeably better. + This is a minimal host. The [`@kaleidorg/mind` README](https://www.npmjs.com/package/@kaleidorg/mind) and the `examples/node-minimal` and `examples/rgb-agent` folders in the [kaleido-mind repository](https://github.com/kaleidoswap/kaleido-mind) go further, with skills, recipes, and a larger model. The example loads Qwen3.5 2B (about 1.3 GB), which handled every signet RGB wallet task in our tests. `QWEN3_5_4B_MULTIMODAL_Q4_K_M` (about 2.7 GB) is as accurate but about twice as slow; the 0.8B model loops on wallet actions. The token caps in the sample leave room for its reasoning before the answer. ## 5. Prompts to Try diff --git a/mintlify-docs/ai-tools/faq.mdx b/mintlify-docs/ai-tools/faq.mdx index 654bc40..594de22 100644 --- a/mintlify-docs/ai-tools/faq.mdx +++ b/mintlify-docs/ai-tools/faq.mdx @@ -41,7 +41,7 @@ description: "Common questions about the KaleidoSwap AI tools, covering custody, It varies by surface: - **KaleidoAgent** reasons with a hosted model, Claude or OpenAI, selected with `AGENT_PROVIDER`. - - **KaleidoMind** runs the model **on the device** through the QVAC SDK, locally or on a desktop you explicitly paired. Its tiered funnel means most requests never reach the model at all. + - **KaleidoMind** runs the model **on the device** through the QVAC SDK. The recommended model is Qwen3.5 2B, on desktop and phone alike. Its tiered funnel means most requests never reach the model at all. Pairing a phone with the desktop app for delegated inference is paused in desktop app v0.5.1. - **MCP servers** are model-agnostic. Whatever model your MCP host runs is the one calling the tools. diff --git a/mintlify-docs/ai-tools/kaleido-mind.mdx b/mintlify-docs/ai-tools/kaleido-mind.mdx index da091cf..17684f7 100644 --- a/mintlify-docs/ai-tools/kaleido-mind.mdx +++ b/mintlify-docs/ai-tools/kaleido-mind.mdx @@ -65,9 +65,23 @@ LLM, embedding, speech-to-text, and text-to-speech inference all run through the | Package | Version | |---------|---------| -| `@qvac/sdk` | The core engine works with 0.13 and later. Install the latest release, which is what the hosts and evals are tested against | +| `@qvac/sdk` | The core engine works with 0.13.1 and later and is tested with 0.21. P2P delegation needs 0.13.1–0.18: QVAC removed it in 0.19 | | `@kaleidorg/mind` | Takes `@qvac/sdk` as an optional peer dependency, so the engine also runs with no model at all against the mock provider | +### Recommended models + +KaleidoMind 0.8 recommends the Qwen3.5 family. Every model below is a `@qvac/sdk` registry constant, so `loadModel({ modelSrc: QWEN3_5_2B_MULTIMODAL_Q4_K_M })` downloads and caches it. Qwen3 and Hermes 3 are no longer recommended: the small Qwen3 models garbled the arguments of multi-field wallet calls such as `rln_issue_asset`. + +| Model | `@qvac/sdk` constant | Download | RAM | Use | +|---|---|---|---|---| +| Qwen3.5 0.8B | `QWEN3_5_0_8B_MULTIMODAL_Q4_K_M` | 0.53 GB | ~1.5 GB | Smoke tests only; loops on wallet actions | +| Qwen3.5 2B | `QWEN3_5_2B_MULTIMODAL_Q4_K_M` | 1.3 GB | ~3 GB | **Default.** Passed all 7 signet RGB wallet tasks, 45–150 s per question on an M4 laptop | +| Qwen3.5 4B | `QWEN3_5_4B_MULTIMODAL_Q4_K_M` | 2.7 GB | ~5 GB | Same correctness as 2B, about twice as slow | +| Qwen3.5 9B | `QWEN3_5_9B_MULTIMODAL_Q4_K_M` | 5.7 GB | ~9 GB | 16 GB machines; slow for interactive chat | +| Qwen3.6 35B-A3B (MoE) | `QWEN3_6_35B_A3B_MULTIMODAL_Q4_K_M` | 22 GB | ~26 GB | 32 GB+ machines | + +The same list is exported as `QWEN35_MODELS` from `@kaleidorg/mind/qvac`. + `createQvacProvider` turns the SDK's `completion` into the engine's `LLMProvider`. The SDK functions are injected rather than imported, so the host owns the model lifecycle (`loadModel`, `unloadModel`) and the engine never pulls QVAC into a bundle that does not need it. ## Build With It @@ -88,7 +102,7 @@ For a step-by-step version on signet, follow [Build a Local RGB Agent](/ai-tools | Host | Role | |------|------| | [Rate](https://github.com/kaleidoswap/Rate) | React Native mobile wallet, with local LLM, speech-to-text, neural text-to-speech, and the hands-free voice loop | -| [Desktop App](/desktop-app/getting-started/introduction) | Runs the engine as in-app chat through the Tauri sidecar (`apps/provider`), which consumes `kaleido-mcp` as a stdio tool source, and can act as the paired inference peer for a phone | +| [Desktop App](/desktop-app/getting-started/introduction) | Runs the engine as in-app chat through the Tauri sidecar (`apps/provider`), which consumes `kaleido-mcp` as a stdio tool source, and can act as the paired inference peer for a phone (phone pairing is paused in desktop app v0.5.1) | The fastest way to try it is the Desktop App. [Download the latest release](https://kaleidoswap.com/downloads) and follow the [installation guide](/desktop-app/getting-started/installation). diff --git a/mintlify-docs/cn/ai-tools/build-local-rgb-agent.mdx b/mintlify-docs/cn/ai-tools/build-local-rgb-agent.mdx index c22d447..6af515b 100644 --- a/mintlify-docs/cn/ai-tools/build-local-rgb-agent.mdx +++ b/mintlify-docs/cn/ai-tools/build-local-rgb-agent.mdx @@ -108,11 +108,11 @@ npm init -y && npm pkg set type=module npm install @kaleidorg/mind @qvac/sdk @modelcontextprotocol/sdk ``` -KaleidoMind 支持 `@qvac/sdk` 0.13 及以上版本,建议安装最新版本。把下面的代码保存为 `agent.mjs`: +KaleidoMind 支持 `@qvac/sdk` 0.13.1 及以上版本,并以 0.21 为测试基准。把下面的代码保存为 `agent.mjs`: ```js agent.mjs import { createInterface } from 'node:readline/promises'; -import { completion, cancel, loadModel, QWEN3_600M_INST_Q4 } from '@qvac/sdk'; +import { completion, cancel, loadModel, QWEN3_5_2B_MULTIMODAL_Q4_K_M } from '@qvac/sdk'; import { Funnel, ToolRegistry, confirmReadback } from '@kaleidorg/mind'; import { McpToolSource } from '@kaleidorg/mind/mcp'; import { createQvacProvider } from '@kaleidorg/mind/qvac'; @@ -136,11 +136,15 @@ await kaleido.connect(); // 2. 通过 QVAC 加载本地模型(首次运行时下载) const modelId = await loadModel({ - modelSrc: QWEN3_600M_INST_Q4, + modelSrc: QWEN3_5_2B_MULTIMODAL_Q4_K_M, modelType: 'llm', modelConfig: { ctx_size: 8192, tools: true }, }); -const provider = createQvacProvider({ completion, cancel, getModelId: () => modelId }); +const provider = createQvacProvider({ + completion, cancel, getModelId: () => modelId, + defaultMaxTokens: 1536, // Qwen3.5 reasons before answering; leave room for both + maxThinkingTokens: 512, +}); // 3. 分层漏斗,任何花费之前都会弹出确认 const funnel = new Funnel({ provider, tools: new ToolRegistry([kaleido]) }); @@ -168,7 +172,7 @@ node agent.mjs 首次启动会下载模型。每个动用资金的工具都会在 `onConfirm` 回调处暂停,在你输入 `y` 之前,任何资金都不会离开节点。 - 这是一个最小的宿主。[`@kaleidorg/mind` README](https://www.npmjs.com/package/@kaleidorg/mind) 以及 [kaleido-mind 仓库](https://github.com/kaleidoswap/kaleido-mind)中的 `examples/node-minimal` 和 `examples/rgb-agent` 目录走得更远,包含 skills、recipe 和更大的模型。0.6B 模型能很好地处理快速路径和 recipe;对于开放式请求,Qwen3 1.7B 或 4B 这类更大的模型表现明显更好。 + 这是一个最小的宿主。[`@kaleidorg/mind` README](https://www.npmjs.com/package/@kaleidorg/mind) 以及 [kaleido-mind 仓库](https://github.com/kaleidoswap/kaleido-mind)中的 `examples/node-minimal` 和 `examples/rgb-agent` 目录走得更远,包含 skills、recipe 和更大的模型。示例加载 Qwen3.5 2B(约 1.3 GB),在我们的测试中它完成了所有 signet RGB 钱包任务。`QWEN3_5_4B_MULTIMODAL_Q4_K_M`(约 2.7 GB)同样准确,但速度约慢一倍;0.8B 模型在钱包操作上会陷入循环。示例中的 token 上限为模型回答前的推理留出了空间。 ## 5. 可以尝试的提示词 {#5-prompts-to-try} diff --git a/mintlify-docs/cn/ai-tools/faq.mdx b/mintlify-docs/cn/ai-tools/faq.mdx index 05bf808..403e22e 100644 --- a/mintlify-docs/cn/ai-tools/faq.mdx +++ b/mintlify-docs/cn/ai-tools/faq.mdx @@ -41,7 +41,7 @@ description: "关于 KaleidoSwap AI 工具的常见问题:托管方式、种 因接入方式而异: - **KaleidoAgent** 用托管模型推理,Claude 或 OpenAI,通过 `AGENT_PROVIDER` 选择。 - - **KaleidoMind** 通过 QVAC SDK 让模型**在设备上**运行,可以在本地,也可以在你显式配对过的桌面端。它的分层漏斗设计意味着大多数请求根本不会到达模型。 + - **KaleidoMind** 通过 QVAC SDK 让模型**在设备上**运行。推荐使用 Qwen3.5 2B,桌面端和手机都适用。它的分层漏斗设计意味着大多数请求根本不会到达模型。桌面应用 v0.5.1 中,手机与桌面配对进行委托推理的功能已暂停。 - **MCP 服务器** 与模型无关。调用工具的就是你的 MCP 宿主所运行的那个模型。 diff --git a/mintlify-docs/cn/ai-tools/kaleido-mind.mdx b/mintlify-docs/cn/ai-tools/kaleido-mind.mdx index c175927..64fe52e 100644 --- a/mintlify-docs/cn/ai-tools/kaleido-mind.mdx +++ b/mintlify-docs/cn/ai-tools/kaleido-mind.mdx @@ -65,9 +65,23 @@ LLM、嵌入、语音转文字和文字转语音推理全部通过 Tether 的本 | 包 | 版本 | |---------|---------| -| `@qvac/sdk` | 核心引擎支持 0.13 及以上版本。建议安装最新版本,宿主与评测都以它为测试基准 | +| `@qvac/sdk` | 核心引擎支持 0.13.1 及以上版本,并以 0.21 为测试基准。P2P 委托需要 0.13.1–0.18:QVAC 在 0.19 中移除了该功能 | | `@kaleidorg/mind` | 把 `@qvac/sdk` 作为可选 peer 依赖,因此引擎在完全没有模型的情况下也能配合 mock provider 运行 | +### 推荐模型 + +KaleidoMind 0.8 推荐 Qwen3.5 系列。下表中的每个模型都是 `@qvac/sdk` 注册表常量,`loadModel({ modelSrc: QWEN3_5_2B_MULTIMODAL_Q4_K_M })` 会自动下载并缓存。不再推荐 Qwen3 和 Hermes 3:较小的 Qwen3 模型会弄乱 `rln_issue_asset` 等多字段钱包调用的参数。 + +| 模型 | `@qvac/sdk` 常量 | 下载大小 | 内存 | 用途 | +|---|---|---|---|---| +| Qwen3.5 0.8B | `QWEN3_5_0_8B_MULTIMODAL_Q4_K_M` | 0.53 GB | 约 1.5 GB | 仅用于冒烟测试;执行钱包操作时会陷入循环 | +| Qwen3.5 2B | `QWEN3_5_2B_MULTIMODAL_Q4_K_M` | 1.3 GB | 约 3 GB | **默认。** 在 M4 笔记本上通过全部 7 项 signet RGB 钱包任务,每个问题 45–150 秒 | +| Qwen3.5 4B | `QWEN3_5_4B_MULTIMODAL_Q4_K_M` | 2.7 GB | 约 5 GB | 正确率与 2B 相同,速度约慢一倍 | +| Qwen3.5 9B | `QWEN3_5_9B_MULTIMODAL_Q4_K_M` | 5.7 GB | 约 9 GB | 16 GB 内存的机器;交互式聊天较慢 | +| Qwen3.6 35B-A3B(MoE) | `QWEN3_6_35B_A3B_MULTIMODAL_Q4_K_M` | 22 GB | 约 26 GB | 32 GB 以上内存的机器 | + +同一份列表也以 `QWEN35_MODELS` 从 `@kaleidorg/mind/qvac` 导出。 + `createQvacProvider` 把 SDK 的 `completion` 转换为引擎所需的 `LLMProvider`。SDK 函数通过注入而非导入的方式传入,因此由宿主掌控模型生命周期(`loadModel`、`unloadModel`),引擎也不会把 QVAC 打进不需要它的包里。 ## 动手构建 {#build-with-it} @@ -88,7 +102,7 @@ LLM、嵌入、语音转文字和文字转语音推理全部通过 Tether 的本 | 宿主 | 角色 | |------|------| | [Rate](https://github.com/kaleidoswap/Rate) | React Native 移动钱包,内含本地 LLM、语音转文字、神经网络文字转语音,以及免手操作的语音循环 | -| [桌面应用](/cn/desktop-app/getting-started/introduction) | 通过 Tauri sidecar(`apps/provider`)把引擎作为应用内聊天运行,该 sidecar 以 stdio 工具源方式接入 `kaleido-mcp`,同时还能作为手机的配对推理节点 | +| [桌面应用](/cn/desktop-app/getting-started/introduction) | 通过 Tauri sidecar(`apps/provider`)把引擎作为应用内聊天运行,该 sidecar 以 stdio 工具源方式接入 `kaleido-mcp`,同时还能作为手机的配对推理节点(桌面应用 v0.5.1 中手机配对功能暂停) | 最快的体验方式是桌面应用。[下载最新版本](https://kaleidoswap.com/downloads),然后按照[安装指南](/cn/desktop-app/getting-started/installation)操作。