diff --git a/plugins/antianqi/mcode-feishu-bridge/.claude-plugin/plugin.json b/plugins/antianqi/mcode-feishu-bridge/.claude-plugin/plugin.json new file mode 100644 index 00000000..23ba7403 --- /dev/null +++ b/plugins/antianqi/mcode-feishu-bridge/.claude-plugin/plugin.json @@ -0,0 +1,25 @@ +{ + "name": "mcode-feishu-bridge", + "version": "0.1.0", + "description": "Drive MiniMax Code remotely from a Lark/Feishu conversation. A message you send in Feishu becomes a local mcode turn with full tool access, and the answer is written back into that same message in place, so a long task does not flood the chat. Requires the lark-cli and mcode CLIs the user already has; ships no credentials of its own.", + "author": { + "name": "antianqi", + "url": "https://github.com/antianqi" + }, + "homepage": "https://github.com/MiniMax-AI/MiniMax-Code-Plugins/tree/main/plugins/antianqi/mcode-feishu-bridge", + "repository": "https://github.com/MiniMax-AI/MiniMax-Code-Plugins.git", + "license": "Apache-2.0", + "keywords": [ + "mcode", + "feishu", + "lark", + "remote-control", + "chatops", + "cli", + "zero-dependency" + ], + "skills": [ + "./skills/SKILL.md", + "./skills/mcode-feishu-bridge/SKILL.md" + ] +} diff --git a/plugins/antianqi/mcode-feishu-bridge/.gitattributes b/plugins/antianqi/mcode-feishu-bridge/.gitattributes new file mode 100644 index 00000000..57b3c7b7 --- /dev/null +++ b/plugins/antianqi/mcode-feishu-bridge/.gitattributes @@ -0,0 +1,11 @@ +# Keep everything in this Plugin byte-stable. +# +# The suite reads this Plugin's own source for two static checks (no shell +# routing, no bare "asc without --page-all" fetch), and a CRLF checkout would +# change what those checks see. JSON must also stay BOM-free: the repository +# validator rejects a BOM in plugin.json. +* text=auto eol=lf +*.json text eol=lf +*.md text eol=lf +*.mjs text eol=lf +LICENSE text eol=lf diff --git a/plugins/antianqi/mcode-feishu-bridge/LICENSE b/plugins/antianqi/mcode-feishu-bridge/LICENSE new file mode 100644 index 00000000..125be1b8 --- /dev/null +++ b/plugins/antianqi/mcode-feishu-bridge/LICENSE @@ -0,0 +1,192 @@ + Apache License + Version 2.0, January 2004 + http://www.apache.org/licenses/ + + TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION + + 1. Definitions. + + "License" shall mean the terms and conditions for use, reproduction, + and distribution as defined by Sections 1 through 9 of this document. + + "Licensor" shall mean the copyright owner or entity authorized by + the copyright owner that is granting the License. + + "Legal Entity" shall mean the union of the acting entity and all + other entities that control, are controlled by, or are under common + control with that entity. For the purposes of this definition, + "control" means (i) the power, direct or indirect, to cause the + direction or management of such entity, whether by contract or + otherwise, or (ii) ownership of fifty percent (50%) or more of the + outstanding shares, or (iii) beneficial ownership of such entity. + + "You" (or "Your") shall mean an individual or Legal Entity + exercising permissions granted by this License. + + "Source" form shall mean the preferred form for making modifications, + including but not limited to software source code, documentation + source, and configuration files. + + "Object" form shall mean any form resulting from mechanical + transformation or translation of a Source form, including but + not limited to compiled object code, generated documentation, + and conversions to other media types. + + "Work" shall mean the work of authorship, whether in Source or + Object form, made available under the License, as indicated by a + copyright notice that is included in or attached to the work + (an example is provided in the Appendix below). + + "Derivative Works" shall mean any work, whether in Source or Object + form, that is based on (or derived from) the Work and for which the + editorial revisions, annotations, elaborations, or other modifications + represent, as a whole, an original work of authorship. For the purposes + of this License, Derivative Works shall not include works that remain + separable from, or merely link (or bind by name) to the interfaces of, + the Work and Derivative Works thereof. + + "Contribution" shall mean any work of authorship, including + the original version of the Work and any modifications or additions + to that Work or Derivative Works thereof, that is intentionally + submitted to Licensor for inclusion in the Work by the copyright owner + or by an individual or Legal Entity authorized to submit on behalf of + the copyright owner. For the purposes of this definition, "submitted" + means any form of electronic, verbal, or written communication sent + to the Licensor or its representatives, including but not limited to + communication on electronic mailing lists, source code control systems, + and issue tracking systems that are managed by, or on behalf of, the + Licensor for the purpose of discussing and improving the Work, but + excluding communication that is conspicuously marked or otherwise + designated in writing by the copyright owner as "Not a Contribution." + + "Contributor" shall mean Licensor and any individual or Legal Entity + on behalf of whom a Contribution has been received by Licensor and + subsequently incorporated within the Work. + + 2. Grant of Copyright License. Subject to the terms and conditions of + this License, each Contributor hereby grants to You a perpetual, + worldwide, non-exclusive, no-charge, royalty-free, irrevocable + copyright license to reproduce, prepare Derivative Works of, + publicly display, publicly perform, sublicense, and distribute the + Work and such Derivative Works in Source or Object form. + + 3. Grant of Patent License. Subject to the terms and conditions of + this License, each Contributor hereby grants to You a perpetual, + worldwide, non-exclusive, no-charge, royalty-free, irrevocable + (except as stated in this section) patent license to make, have made, + use, offer to sell, sell, import, and otherwise transfer the Work, + where such license applies only to those patent claims licensable + by such Contributor that are necessarily infringed by their + Contribution(s) alone or by combination of their Contribution(s) + with the Work to which such Contribution(s) was submitted. If You + institute patent litigation against any entity (including a + cross-claim or counterclaim in a lawsuit) alleging that the Work + or a Contribution incorporated within the Work constitutes direct + or contributory patent infringement, then any patent licenses + granted to You under this License for that Work shall terminate + as of the date such litigation is filed. + + 4. Redistribution. You may reproduce and distribute copies of the + Work or Derivative Works thereof in any medium, with or without + modifications, and in Source or Object form, provided that You + meet the following conditions: + + (a) You must give any other recipients of the Work or + Derivative Works a copy of this License; and + + (b) You must cause any modified files to carry prominent notices + stating that You changed the files; and + + (c) You must retain, in the Source form of any Derivative Works + that You distribute, all copyright, patent, trademark, and + attribution notices from the Source form of the Work, + excluding those notices that do not pertain to any part of + the Derivative Works; and + + (d) If the Work includes a "NOTICE" text file as part of its + distribution, then any Derivative Works that You distribute must + include a readable copy of the attribution notices contained + within such NOTICE file, excluding those notices that do not + pertain to any part of the Derivative Works, in at least one + of the following places: within a NOTICE text file distributed + as part of the Derivative Works; within the Source form or + documentation, if provided along with the Derivative Works; or, + within a display generated by the Derivative Works, if and + wherever such third-party notices normally appear. The contents + of the NOTICE file are for informational purposes only and + do not modify the License. You may add Your own attribution + notices within Derivative Works that You distribute, alongside + or as an addendum to the NOTICE text from the Work, provided + that such additional attribution notices cannot be construed + as modifying the License. + + You may add Your own copyright statement to Your modifications and + may provide additional or different license terms and conditions + for use, reproduction, or distribution of Your modifications, or + for any such Derivative Works as a whole, provided Your use, + reproduction, and distribution of the Work otherwise complies with + the conditions stated in this License. + + 5. Submission of Contributions. Unless You explicitly state otherwise, + any Contribution intentionally submitted for inclusion in the Work + by You to the Licensor shall be under the terms and conditions of + this License, without any additional terms or conditions. + Notwithstanding the above, nothing herein shall supersede or modify + the terms of any separate license agreement you may have executed + with Licensor regarding such Contributions. + + 6. Trademarks. This License does not grant permission to use the trade + names, trademarks, service marks, or product names of the Licensor, + except as required for reasonable and customary use in describing the + origin of the Work and reproducing the content of the NOTICE file. + + 7. Disclaimer of Warranty. Unless required by applicable law or + agreed to in writing, Licensor provides the Work (and each + Contributor provides its Contributions) on an "AS IS" BASIS, + WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or + implied, including, without limitation, any warranties or conditions + of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A + PARTICULAR PURPOSE. You are solely responsible for determining the + appropriateness of using or redistributing the Work and assume any + risks associated with Your exercise of permissions under this License. + + 8. Limitation of Liability. In no event and under no legal theory, + whether in tort (including negligence), contract, or otherwise, + unless required by applicable law (such as deliberate and grossly + negligent acts) or agreed to in writing, shall any Contributor be + liable to You for damages, including any direct, indirect, special, + incidental, or consequential damages of any character arising as a + result of this License or out of the use or inability to use the + Work (including but not limited to damages for loss of goodwill, + work stoppage, computer failure or malfunction, or any and all + other commercial damages or losses), even if such Contributor + has been advised of the possibility of such damages. + + 9. Accepting Warranty or Additional Liability. While redistributing + the Work or Derivative Works thereof, You may choose to offer, + and charge a fee for, acceptance of support, warranty, indemnity, + or other liability obligations and/or rights consistent with this + License. However, in accepting such obligations, You may act only + on Your own behalf and on Your sole responsibility, not on behalf + of any other Contributor, and only if You agree to indemnify, + defend, and hold each Contributor harmless for any liability + incurred by, or claims asserted against, such Contributor by reason + of your accepting any such warranty or additional liability. + + END OF TERMS AND CONDITIONS + + APPENDIX: How to apply the Apache License to your work. + + Copyright 2026 MCode Plugins contributors + + Licensed under the Apache License, Version 2.0 (the "License"); + you may not use this file except in compliance with the License. + You may obtain a copy of the License at + + http://www.apache.org/licenses/LICENSE-2.0 + + Unless required by applicable law or agreed to in writing, software + distributed under the License is distributed on an "AS IS" BASIS, + WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. + See the License for the specific language governing permissions and + limitations under the License. diff --git a/plugins/antianqi/mcode-feishu-bridge/README.md b/plugins/antianqi/mcode-feishu-bridge/README.md new file mode 100644 index 00000000..2ba5f29a --- /dev/null +++ b/plugins/antianqi/mcode-feishu-bridge/README.md @@ -0,0 +1,452 @@ +# mcode-feishu-bridge + +Drive [MiniMax Code](https://github.com/MiniMax-AI/MiniMax-Code) remotely from a +Lark/Feishu conversation. Send a task from your phone, mcode runs it on your machine +with its normal tool access, and the answer comes back **in the same message**. + +- **The problem.** MiniMax Code is a terminal application. On a desktop that is + fine; away from it, the only options are a remote desktop session or a tunnel. + Neither is pleasant, and neither works from a phone. + +- **What this does.** A small polling daemon watches one Feishu conversation. Each + new message becomes one `mcode exec` turn in a workspace directory that belongs to + that conversation. The answer is written back into the same message by editing it, + so a task that makes twelve tool calls still produces one chat bubble. + +## Example + +**Prompt** (sent in the Feishu conversation): + +> create a file notes.md in the current directory with the content "shopping list", +> then read it back to confirm + +**Result** — the bot's reply in the chat, edited in place: + +> 🧠 received, thinking… +> +> ⚙️ calling `write`… +> +> Created and read back successfully. +> +> `…\workspaces\oc_xxxxxx\notes.md` — 16 bytes, content: +> +> ``` +> shopping list +> ``` +> +> `shopping list` = 13 characters + newline = 14 bytes, consistent with the byte count +> and the content read back matches character for character. +> +> --- +> ⏱ 11.4s · 2 tool calls + +The intermediate states and the final answer are the same Feishu message, so the +conversation stays readable and nothing is duplicated. + +## What you get + +- **One message per task.** The placeholder, every tool call, and the final answer + are the same Feishu message, edited in place. A turn that runs twelve tools + produces one bubble. +- **Live progress.** The message flips to `⚙️ calling \`write\`…` while a tool runs, + so a two-minute task does not look like a hang. +- **Real tool access.** mcode runs with its normal tools in a per-conversation + workspace, so the bridge can create files, run commands, and read results. +- **Continuous context.** The conversation keeps one mcode session, so "and now + read it back" works. Each conversation gets its own session, so two do not bleed + into each other. +- **Attachments.** Images and files you send are downloaded into the conversation's + workspace and passed to mcode. +- **A hung turn is killed, not waited on.** After a configurable ceiling the process + tree is terminated and the chat gets an explicit notice naming how many tool calls + had already run. +- **A footer on every answer** — elapsed time, token count, tool-call count. +- **Delivery that does not lose your result.** If the edit fails, the answer is sent + as a reply instead. If that fails too, it is retried with a linear backoff and + then given up on loudly, rather than retried forever. +- **One bridge at a time.** A second watcher is refused rather than silently racing + the first one over the same state. +- **Readable while running.** The log is written by the bridge itself, so you can + read it without the file being locked. + +## How this differs from the built-in Feishu support + +MiniMax Code ships Feishu support in two places already, and neither one is this +Plugin. This one sends messages in the opposite direction, against a local mcode +CLI rather than the desktop runtime. + +| | direction | what it is for | +|---|---|---| +| `lark-tools` (built-in Skill) | mcode → Feishu | office work: read and write docs, calendar, Base, mail, approvals | +| Feishu channel (local-runtime) | Feishu → agent | driving the Mavis desktop runtime through a bound Feishu app | +| **This Plugin** | **Feishu → mcode** | **driving a local mcode CLI install** | + +The first is not a remote control. Ask it to summarise this week's meeting notes +and it does that inside Feishu; it does not run a task on your machine. + +The second has the same shape as this Plugin but a different host. It lives in the +Mavis desktop runtime, so it needs the desktop app running and a Feishu app bound +to it, and it answers with interactive cards — including a permission card that +approves a pending tool call by tapping. Prefer it if the desktop app is already +your working environment. + +This Plugin covers the remaining case: a headless machine with `mcode` and +`lark-cli` on `PATH` and nothing else. It is a plain polling process, so any +supervisor you already use can restart it, and there is nothing to install. In +exchange it edits plain text messages instead of sending cards, and it runs with +`--permission full` because there is no in-chat approval prompt to tap — see +[Security model](#security-model). + +The two do not conflict. Installing this Plugin neither requires nor disables the +built-in channel, and both can run at once. + +## Setup + +The Plugin is Skill-only: no `mcp.json`, no `package.json`, no install step. But it +does need three things lined up, and two of them are easy to get half-right. + +### 1. MiniMax Code + +Installed and runnable: + +```bash +mcode --version +``` + +### 2. A Feishu/Lark app, with **two** working identities + +This is the part that trips people up. The bridge reads and writes through +**different identities**, and both have to work: + +| what | identity | why | +|---|---|---| +| read the conversation, download attachments | **user** | the bot cannot see a p2p conversation's history | +| send, reply, and edit the answer | **bot** | only a bot may edit a message it sent | + +```bash +lark-cli auth status +``` + +Both `identities.bot.status` and `identities.user.status` must read `ready`. +If either is not, fix that before going further — otherwise the bridge reads +nothing or writes nothing, and the log only says `fetch messages failed`. + +The app needs at least these scopes, derived from the operations it performs: + +| scope | needed for | identity | +|---|---|---| +| `im:message:readonly` | reading messages in the conversation | user | +| `im:resource` | downloading an image or file you send | user | +| `im:message` | sending the placeholder and the answer | bot | +| `im:message:update` | editing that placeholder into the answer | bot | +| `offline_access` | letting the user token refresh instead of expiring | user | + +`offline_access` is the one people miss. Without it the user token stops working +after a couple of hours and the bridge goes quiet, with no error at the time it +happens. Note also that the user identity here is **your** account: the bridge acts +as you, with everything that implies. + +Publish the app (Feishu requires an app to be published before another user can +talk to it), then open a chat with it. A p2p chat with your own app is the simplest +arrangement and the one this Plugin is built and tested for. + +### 3. A chat id + +There is deliberately no default: a chat id is a private identifier, so it has to be +yours to supply. + +```bash +lark-cli im +chat-list --types=p2p,group --page-all --as user --format json +``` + +Find your conversation in the output. The `--page-all` matters — without it the +listing silently truncates and you may not see the conversation you are looking for. + +### 4. Start it + +```bash +node scripts/mcode-feishu-bridge.mjs --watch --chat +``` + +On the first run the bridge creates a workspace for that conversation and picks up +nothing historical: it starts from the newest message at the time it launches. Send +it a message after starting, not before. + +Full operating instructions are in +[`skills/mcode-feishu-bridge/SKILL.md`](skills/mcode-feishu-bridge/SKILL.md). + +## Usage + +```bash +# one pass, then exit — the quickest way to try it +node scripts/mcode-feishu-bridge.mjs --chat + +# the long-running watcher +node scripts/mcode-feishu-bridge.mjs --watch --chat + +# stop the running watcher; kills any mcode turn with it +node scripts/mcode-feishu-bridge.mjs --stop +``` + +| flag | default | meaning | +|---|---|---| +| `--chat ` | none | conversation to watch; required | +| `--watch` | off | keep polling instead of one pass | +| `--interval ` | 3000 | poll interval | +| `--timeout ` | 600000 | hard ceiling on one mcode turn | +| `--log ` | `/bridge.log` in watch mode | also append output to this file | +| `--stop` | — | stop the running watcher | + +Only `--watch` takes the single-instance lock, so one-shot runs never block each +other. + +The conversation can also be configured once, instead of on every start: + +```bash +MCODE_FEISHU_CHAT=oc_xxxxxxxxxxxxxxxx # environment +# or /config.json -> { "chatId": "oc_xxxxxxxxxxxxxxxx" } +``` + +## What's in the package + +```text +mcode-feishu-bridge/ +├── plugin.json portable Agent Plugins 1.0 manifest +├── .claude-plugin/plugin.json v0.4.0+ manifest +├── skills/ +│ ├── SKILL.md v0.4.0+ Skill +│ └── mcode-feishu-bridge/ +│ └── SKILL.md v0.3.x Skill, byte-identical to the above +├── scripts/ +│ ├── mcode-feishu-bridge.mjs the bridge +│ └── mcode-feishu-bridge.test.mjs the suite +├── README.md +└── LICENSE +``` + +Two Skill copies because the two runtimes discover Skills differently; the +recommended cross-version layout keeps them byte-identical, and this Plugin follows +it. There is no `mcp.json`, no `package.json`, and no bundled binary. + +## Running it in the background + +The Plugin ships no service manager, on purpose: a login hook that runs a hidden +process forever is exactly the kind of thing that should be a deliberate, visible +choice by the person running the machine, not a default a Plugin installs. + +If you do want it always on, run the watcher under whatever supervisor you already +use, and make sure it inherits a real `PATH` so the bridge can find `lark-cli` and +`mcode`: + +```bash +nohup node scripts/mcode-feishu-bridge.mjs --watch --chat >> bridge.log 2>&1 & +``` + +Two rules that follow from how this is built: + +- **Do not redirect the watcher's stdout from a supervisor that waits on the + child's pipe.** A long-lived child spawned with inherited std handles keeps that + pipe open, and the supervisor blocks forever. Start the watcher with std handles + closed or detached, or let the bridge write its own log via `--log`. +- **The bridge takes a single-instance lock, so a second copy is refused with exit + code 3** rather than corrupting state. If that is not what you wanted, find the + first one: `bridge.lock` in the data directory holds its pid. + +## Uninstall + +There is nothing to uninstall. To stop it and remove its data: + +```bash +node scripts/mcode-feishu-bridge.mjs --stop +``` + +Then delete the data directory (`$PLUGIN_DATA`, else `~/.mcode-feishu-bridge`). +That removes the per-conversation workspaces and the downloaded media. Removing the +Plugin from MiniMax Code leaves that directory alone, by design — it is the user's +data, and a Plugin should not delete it on uninstall. + +## Network + +Outbound HTTPS to the Feishu/Lark Open Platform, through `lark-cli`. No inbound +listener, no local port, no other host. If your machine reaches Feishu through a +proxy, `lark-cli` needs it configured; the bridge does not set one. + +## How it works + +``` +Feishu message + │ + ├─ poll (default every 3s) → +chat-messages-list + │ newest page first; full pagination only when the watermark fell off the page + │ + ├─ send a placeholder, then edit it for every state change + │ edits are queued serially, so a late tool hint can never overwrite the answer + │ + ├─ mcode exec --cwd --session --permission full + │ stream-json is parsed as it arrives to show tool calls live + │ + └─ the answer overwrites the placeholder, with a time/token/tool-call footer +``` + +Each conversation gets its own workspace and its own mcode session, so context is +continuous per chat and isolated between chats. + +## What it deliberately does not do + +- **No conversation history replay.** The bridge starts from the last message it + handled. It does not re-run old messages. +- **No concurrency per chat.** The second message waits for the first to finish. + Two `mcode exec` runs against one session would interleave. +- **No card messages.** Feishu cards need a different update path; this uses plain + text messages and edits them. +- **No message deletion.** It never recalls anything. Every message the bot sends + stays until you delete it yourself. + +## Security model + +**This is the part to read before enabling it.** + +- mcode runs with `--permission full`, so it does not stop to ask for approval. A + Feishu message can therefore cause arbitrary changes to your working directory: + writing files, running commands, installing things. A phone is a bad place to + answer approval prompts, so the prompts are removed — but that means the + guarantee you have to rely on is the Feishu account, not the permission system. +- If that trade is not acceptable for you, point the bridge at a dedicated + workspace directory, so the blast radius is one directory you can throw away. + The workspace is created under `/workspaces//`. +- The account used to read the chat is a **user** token, so the bridge acts as you, + with everything that implies. + +## Disclosures + +### No credentials of its own + +This Plugin contains no tokens, keys, secrets, cookies, or private endpoints, and it +reads none from the environment to authenticate anything. There is no credential to +leak from the package. + +### It depends on a lark-cli configuration you already own + +The bridge never signs in. It shells out to `lark-cli`, which you install and +authenticate yourself with `lark-cli config init` and `lark-cli auth login`. Your +`lark-cli` credentials live in your own `lark-cli` config directory and are used by +`lark-cli`, not by this Plugin. If `lark-cli` is not signed in, the bridge reports the +failure and stops; it will not attempt to authenticate on your behalf. + +### No telemetry + +The Plugin collects nothing and sends nothing anywhere except the Feishu Open +Platform, as described below. There is no analytics, no crash reporting, no usage +counter, no phone-home, and no update check. The only thing written to disk is +described under Data. + +### Third-party service: Lark/Feishu + +The only third party involved is the Lark/Feishu Open Platform, operated by ByteDance. +The bridge talks to it exclusively through your own `lark-cli`. It contacts: + +| destination | why | data sent | +|---|---|---| +| `open.feishu.cn` (or `open.larksuite.com`) | read messages, send a message, edit a message, download an attachment | the message text of the conversation being watched, and the text the bridge writes back | +| the Lark/Feishu Open Platform auth host | token refresh, performed by `lark-cli` | whatever `lark-cli` sends; this Plugin does not read or store it | + +Your Feishu organization applies its own retention and access policies to those +messages. Sending a task to this bridge means sending it to Feishu. + +Nothing else is contacted: no registry, no CDN, no analytics endpoint, no GitHub, no +model provider beyond the MiniMax Code installation you already run locally. + +## Data written to disk + +All under the data directory, which is `$PLUGIN_DATA` when the runtime provides one, +otherwise `$MCODE_FEISHU_BRIDGE_DATA`, otherwise `~/.mcode-feishu-bridge`: + +| path | contents | +|---|---| +| `state.json` | per chat: workspace path, mcode session id, last handled message id, turn count, and at most one pending delivery | +| `config.json` | only what you put in it; the bridge reads `chatId` and nothing else | +| `bridge.log` | the last run's log lines, rotated to `bridge.log.1` on each start | +| `bridge.lock` | the running pid | +| `workspaces//` | the mcode working directory for that chat | +| `media/` | images and files downloaded from the chat, so mcode can read them | + +`state.json` is written atomically (staged, fsynced, renamed), so an interrupted +write cannot leave a truncated file that would make the bridge reprocess the whole +conversation. + +## Tests + +```bash +node scripts/mcode-feishu-bridge.test.mjs +``` + +84 assertions. They import the real module rather than a copy, so a green run is +evidence about the shipped code. Coverage: + +- edits to one message are serialised, so a slow tool hint cannot overwrite the answer +- a failed edit falls back to a reply instead of stranding the user on a placeholder +- delivery failure retries with a linear backoff and then gives up, never looping +- the watermark always advances, so a persistent failure cannot re-run mcode +- a hung mcode turn is killed as a process tree, and resolution waits for the process + to actually exit +- one watcher at a time; a stale lock is reclaimed +- message paging and ordering, including a regression test built from a real + conversation that had outgrown one page +- fetch problems throw rather than returning an empty list +- the source never routes a child process through a shell, and never sends `--text` + +Two of these are regression tests for defects that shipped in an earlier draft and +were found in real use. Their comments record what actually happened, including the +measured numbers. + +## Traps this Plugin works around + +Each of these was reproduced, not anticipated. The code comments say the same thing +next to the fix. + +1. **Never forward arguments through `cmd /c`.** Node builds a command line and + `cmd.exe` parses it again. mcode answers can contain + `` markup, and the `<`, `>` and `"` inside it are + redirection and quote operators to `cmd`. The command is shredded: exit 1, empty + stdout, empty stderr, no diagnosable cause, and a retry loop that floods the chat. + Measured with one payload against one message: `cmd /c` → exit 1, no output; + direct spawn → exit 0, delivered. Every child process here is spawned directly + with `shell: false`. +2. **Message paging and ordering.** `--order asc --page-size 50` returns the + *oldest* 50 messages. Once a conversation outgrows one page, new messages are + invisible and the bridge looks alive but deaf, logging nothing. It now fetches the + newest page and escalates to full pagination only when the watermark is not on it. +3. **Fetch problems must not be silent.** Returning a status code and trusting + someone to check it is not a safeguard: deleting that one line left the whole + suite green. `collectFresh` throws instead, so the only exit is a catch that + already logs. +4. **A timeout has to wait for the process to die.** `taskkill` is asynchronous. + Resolving immediately let the next message race a zombie mcode for the same + workspace and session. +5. **Windows PowerShell 5.1 reads BOM-less UTF-8 as the local code page.** Relevant + if you write a launcher script for it; not applicable to this Plugin, which ships + no `.ps1`. + +## When it does not work + +| symptom | most likely cause | check | +|---|---|---| +| no reply at all, log says `fetch messages failed` | the **user** identity is not usable: token expired, or `offline_access` was never granted | `lark-cli auth status` — look at `identities.user` | +| the bridge never starts, exits 1 at launch | `lark-cli` or `mcode` is not on the `PATH` the bridge inherited | run `which lark-cli` / `which mcode` in that same environment | +| a reply appears, but it is always the failure notice | the **bot** identity cannot send, so even the fallback reply fails | `lark-cli auth status` — look at `identities.bot` | +| works for hours, then goes quiet | the user token expired because `offline_access` was not granted at authorisation time | re-authorise with `lark-cli auth login` and request that scope | +| attachments are ignored | missing `im:resource` | the app's permission list | +| nothing happens, and the log says the watermark is not on the page | the bridge fell too far behind and the full-pagination fallback also failed | restart the watcher; it re-reads the newest page | +| a second watcher refuses to start with exit 3 | one is already running, which is the point | `bridge.lock` holds its pid; use `--stop` | +| a turn is reported as timed out | mcode hung, most often on an interactive prompt | the notice names how many tool calls had already run; check the workspace | + +None of these are silent: every one of them produces a line in `bridge.log`. If the +log is empty and the chat is silent, the watcher is not running at all. + +The same table, in the order to work through it when a user reports silence, is in +[`skills/mcode-feishu-bridge/SKILL.md`](skills/mcode-feishu-bridge/SKILL.md), along +with the identity and scope pre-flight an agent should check before starting. + +## License + +Apache-2.0. See [LICENSE](LICENSE). diff --git a/plugins/antianqi/mcode-feishu-bridge/plugin.json b/plugins/antianqi/mcode-feishu-bridge/plugin.json new file mode 100644 index 00000000..a0f9a764 --- /dev/null +++ b/plugins/antianqi/mcode-feishu-bridge/plugin.json @@ -0,0 +1,22 @@ +{ + "$schema": "https://agent-plugins.org/schemas/1.0.0/plugin.schema.json", + "name": "mcode-feishu-bridge", + "version": "0.1.0", + "description": "Drive MiniMax Code remotely from a Lark/Feishu conversation. A message you send in Feishu becomes a local mcode turn with full tool access, and the answer is written back into that same message in place, so a long task does not flood the chat. Requires the lark-cli and mcode CLIs the user already has; ships no credentials of its own.", + "author": { + "name": "antianqi", + "url": "https://github.com/antianqi" + }, + "homepage": "https://github.com/MiniMax-AI/MiniMax-Code-Plugins/tree/main/plugins/antianqi/mcode-feishu-bridge", + "repository": "https://github.com/MiniMax-AI/MiniMax-Code-Plugins.git", + "license": "Apache-2.0", + "keywords": [ + "mcode", + "feishu", + "lark", + "remote-control", + "chatops", + "cli", + "zero-dependency" + ] +} diff --git a/plugins/antianqi/mcode-feishu-bridge/scripts/mcode-feishu-bridge.mjs b/plugins/antianqi/mcode-feishu-bridge/scripts/mcode-feishu-bridge.mjs new file mode 100644 index 00000000..adc4d613 --- /dev/null +++ b/plugins/antianqi/mcode-feishu-bridge/scripts/mcode-feishu-bridge.mjs @@ -0,0 +1,1076 @@ +/** + * mcode-feishu-bridge + * + * Drive MiniMax Code remotely from a Lark/Feishu conversation: a message in + * Feishu becomes a local mcode turn, and the answer is written back into that + * same message in place. + * + * Zero npm dependencies. It shells out to two CLIs the user already has: + * `lark-cli` and `mcode`. It ships no credentials of its own. + * + * Cross-platform traps (all reproduced on Windows, not theoretical): + * + * 1. [The expensive one] Never forward arguments through `cmd /c`. + * Node builds a command line, then cmd.exe parses it again — two parses. + * mcode answers can contain resource markup such as + * , and the < > " inside it are + * redirection and quote operators to cmd. The command is shredded: + * exit 1, empty stdout, empty stderr, no diagnosable cause, and the + * failure then drives a retry loop that floods the chat. + * Measured with the same payload against the same message: + * cmd /c lark-cli ... -> exit 1, no output + * spawn(lark-cli) -> exit 0, delivered + * So every child process is spawned directly, with shell:false, and the + * shell is never involved. + * + * 2. `>` and `2>` are not redirections when passed as argv. A binary does not + * parse them; it treats them as ordinary arguments and fails. + * + * 3. lark-cli implements the `+xxx` subcommands in bin/lark-cli (a compiled + * binary); scripts/run.js only forwards to it. When bypassing the shell, + * spawn the binary, not run.js. + * + * 4. lark-cli does not paginate +chat-list or +chat-messages-list by default. + * See selectFresh(): ordering and paging are what make the bridge go deaf + * when a conversation outgrows one page. + * + * 5. --text silently drops text containing a newline. Always send --content + * with a JSON body. + */ +import { spawn, spawnSync } from "node:child_process"; +import { readFileSync, writeFileSync, appendFileSync, mkdirSync, existsSync, unlinkSync, + readdirSync, renameSync, openSync, closeSync, fsyncSync } from "node:fs"; +import { join, dirname } from "node:path"; +import { tmpdir, homedir } from "node:os"; +import { inspect } from "node:util"; + +const IS_WIN = process.platform === "win32"; + +// ─────────────────────────── configuration ─────────────────────────── +/** + * Data directory. The runtime reserves PLUGIN_DATA for a Plugin; honour it so + * the bridge is sandboxed with the rest of the Plugin, and fall back to a + * per-user directory rather than the shared temp dir (a temp dir is wiped and + * is world-writable, neither of which suits state and downloaded attachments). + */ +const DATA_DIR = process.env.PLUGIN_DATA + || process.env.MCODE_FEISHU_BRIDGE_DATA + || join(homedir(), ".mcode-feishu-bridge"); +const STATE_FILE = join(DATA_DIR, "state.json"); +const LOCK_FILE = join(DATA_DIR, "bridge.lock"); +const CONFIG_FILE = join(DATA_DIR, "config.json"); +const LOG_FILE_DEFAULT = join(DATA_DIR, "bridge.log"); +const WS_ROOT = join(DATA_DIR, "workspaces"); +const DOWNLOAD_ROOT = join(DATA_DIR, "media"); +const TMP = tmpdir(); + +const PERMISSION = "full"; + +const args = process.argv.slice(2); +/** + * The chat to watch. Never ship a default: a chat id is a private identifier + * and belongs to whoever configured it. + * Resolution order: --chat, then $MCODE_FEISHU_CHAT, then config.json. + */ +const CHAT_ID = argOf("--chat") || process.env.MCODE_FEISHU_CHAT || configChatId(); +const WATCH = args.includes("--watch"); +const INTERVAL = Number(argOf("--interval")) || 3000; +/** Hard ceiling on one mcode turn. The default is 10 minutes: a normal turn + * takes 10-15s, so 10 minutes is generous while still short enough that a + * hung agent does not leave the chat waiting forever. */ +const MCODE_TIMEOUT_MS = Number(argOf("--timeout")) || 10 * 60 * 1000; +/** Grace period after a kill is issued, waiting for the process to really die. */ +const KILL_GRACE_MS = 3000; + +function argOf(flag) { + const i = args.indexOf(flag); + return i >= 0 ? args[i + 1] : null; +} + +/** Read a non-secret setting from the data-dir config file. */ +function configValue(key) { + try { + const raw = JSON.parse(readFileSync(CONFIG_FILE, "utf8")); + const v = raw?.[key]; + return typeof v === "string" && v.trim() ? v.trim() : null; + } catch { return null; } +} +function configChatId() { return configValue("chatId"); } + +// ─────────────────────────── self-logging ─────────────────────────── +/** + * --log (default: /bridge.log) appends console output to a file. + * + * Redirecting from the launcher instead looks simpler, but a long-lived child + * spawned that way still inherits the launcher's std handles, so an automated + * caller (a login hook, for example) blocks until the child's pipe closes. + * Having the process write its own log keeps it fully detached, and the log + * stays readable while the bridge runs. + */ +const LOG_FILE = argOf("--log") ?? (WATCH ? LOG_FILE_DEFAULT : null); +if (LOG_FILE) { + mkdirSync(dirname(LOG_FILE), { recursive: true }); + for (const k of ["log", "warn", "error"]) { + const orig = console[k].bind(console); + console[k] = (...a) => { + orig(...a); + try { appendFileSync(LOG_FILE, a.map((x) => (typeof x === "string" ? x : inspect(x))).join(" ") + "\n", "utf8"); } + catch {} + }; + } +} + +// ─────────────────────────── executable resolution ─────────────────────────── +/** Find a command on PATH. Cross-platform; no drive letter or absolute prefix is hardcoded. */ +function resolveOnPath(name) { + try { + const r = spawnSync(IS_WIN ? "where.exe" : "which", [name], { + encoding: "utf8", windowsHide: true, + }); + return String(r.stdout || "").split(/\r?\n/).map((s) => s.trim()).filter(Boolean)[0] || null; + } catch { return null; } +} + +/** + * Locate the real lark-cli binary. + * + * The npm shim on Windows is `/lark-cli.cmd`, which runs node run.js, + * and run.js in turn execFileSync's `/node_modules/@larksuite/cli/bin/lark-cli.exe`. + * So derive the binary from the shim found on PATH; never hardcode a drive letter. + */ +function resolveLarkExe() { + if (process.env.MCODE_FEISHU_LARK_BIN && existsSync(process.env.MCODE_FEISHU_LARK_BIN)) { + return process.env.MCODE_FEISHU_LARK_BIN; + } + const shim = (IS_WIN && resolveOnPath("lark-cli.cmd")) || resolveOnPath("lark-cli"); + if (shim) { + const binName = IS_WIN ? "lark-cli.exe" : "lark-cli"; + const cand = join(dirname(shim), "node_modules", "@larksuite", "cli", "bin", binName); + if (existsSync(cand)) return cand; + } + // Some installs expose the binary itself on PATH. + return resolveOnPath(IS_WIN ? "lark-cli.exe" : "lark-cli"); +} + +/** + * Locate the mcode JS entry point. + * + * The Windows launcher is `node.exe cli.js %*`, i.e. plain JavaScript, so it can + * be run by spawning node directly and the shell is bypassed entirely. The + * launcher directory is discovered from the `mcode` command on PATH, and the + * active release is read from the install's own `current` marker file. + */ +function resolveMcodeCli() { + if (process.env.MCODE_FEISHU_MCODE_CLI && existsSync(process.env.MCODE_FEISHU_MCODE_CLI)) { + return process.env.MCODE_FEISHU_MCODE_CLI; + } + const cliAt = (base) => { + const rel = (() => { try { return readFileSync(join(base, "current"), "utf8").trim(); } catch { return null; } })(); + const tryVer = (v) => (v ? join(base, "releases", v, "node_modules", "@minimax-ai", "code", "cli.js") : null); + if (rel && existsSync(tryVer(rel))) return tryVer(rel); + // `current` is missing or points at a version that is not installed: + // fall back to any release directory, newest first. + try { + for (const v of readdirSync(join(base, "releases")).sort().reverse()) { + if (existsSync(tryVer(v))) return tryVer(v); + } + } catch {} + return null; + }; + + // Install roots, discovered from PATH where possible so nothing is hardcoded. + const roots = []; + const push = (p) => { if (p && !roots.includes(p)) roots.push(p); }; + push(join(homedir(), ".minimax-code")); + const shim = resolveOnPath("mcode.cmd") || resolveOnPath("mcode"); + if (shim) push(dirname(dirname(shim))); + if (process.env.MCODE_HOME) push(process.env.MCODE_HOME); + + for (const r of roots) { + const hit = cliAt(r); + if (hit) return hit; + } + return null; +} + +const LARK_EXE = resolveLarkExe(); +const MCODE_CLI = resolveMcodeCli(); + +// Fail fast when a required executable is missing — but not under the test +// harness. The suite imports this module on hosts (CI included) that have no +// mcode installed, and an exit(1) at import time would fail the whole file +// before a single assertion runs. The suite probes resolveMcodeCli() itself and +// skips the live-host cases. +if ((!LARK_EXE || !MCODE_CLI) && process.env.MCODE_FEISHU_BRIDGE_TEST !== "1") { + console.error("Cannot locate the required executables:"); + if (!LARK_EXE) console.error(" lark-cli binary (override with MCODE_FEISHU_LARK_BIN)"); + if (!MCODE_CLI) console.error(" mcode cli.js (override with MCODE_FEISHU_MCODE_CLI)"); + process.exit(1); +} + +// ─────────────────────────── exec layer ─────────────────────────── +/** + * Run an executable, collecting stdout/stderr over pipes, never through a shell. + * + * Why not `cmd /c` plus `> file`: see notes 1 and 2 in the file header. + * Buffer.toString("utf8") is also encoding-safe; a console code page affects + * terminal display, not Buffer decoding. + * + * Returns {code, text, err, ok}. `ok` only means the process exited cleanly and + * says nothing about success: lark-cli can exit 0 on a Feishu-side failure, so + * the caller must inspect the payload. + */ +function runCli(exe, argv, { onData } = {}) { + return new Promise((resolve) => { + const p = spawn(exe, argv, { + shell: false, + windowsHide: true, + stdio: ["ignore", "pipe", "pipe"], + }); + let text = "", err = "", settled = false; + const done = (code) => { + if (settled) return; + settled = true; + resolve({ code, text, err, ok: code === 0 }); + }; + p.stdout.on("data", (d) => { + const s = d.toString("utf8"); + text += s; + onData?.(s); + }); + p.stderr.on("data", (d) => (err += d.toString("utf8"))); + p.on("close", done); + p.on("error", (e) => { err += String(e); done(-1); }); + }); +} + +const lark = (argv) => runCli(LARK_EXE, argv); +const larkJson = async (argv) => { + const { text } = await lark([...argv, "--format", "json"]); + try { return JSON.parse(text); } catch { return null; } +}; + +// ─────────────────────────── mcode layer ─────────────────────────── +/** + * Kill a whole process tree. + * + * child.kill() only reaps the direct child. mcode spawns its own children + * (tools, plugin hosts) which keep file handles and locks if left alive, so + * Windows needs taskkill /T for a tree kill, and POSIX needs a process-group + * signal. Both are invoked without a shell. + */ +function killTree(pid) { + if (!pid) return; + if (IS_WIN) { + try { + const p = spawn("taskkill.exe", ["/F", "/T", "/PID", String(pid)], + { shell: false, windowsHide: true, stdio: "ignore" }); + p.on("error", () => { try { process.kill(pid, "SIGKILL"); } catch {} }); + } catch { + try { process.kill(pid, "SIGKILL"); } catch {} + } + return; + } + // POSIX: signal the group when we can, otherwise the single process. + try { process.kill(-pid, "SIGKILL"); } catch { + try { process.kill(pid, "SIGKILL"); } catch {} + } +} + +/** + * Run one mcode turn. Uses stream-json so a status message can be sent at key + * moments (a tool call, for example). + * + * Spawns node on cli.js directly and consumes stdout line by line through a + * pipe — far simpler than the old "redirect to a file, then poll the file size + * with setInterval", and genuinely streaming. That also disposes of an old + * problem: the file-polling dance existed to dodge PowerShell's GBK, but + * Buffer.toString("utf8") never had an encoding problem, so it was all + * self-inflicted. + * + * onProgress(kind, detail): + * kind = "tool" a tool call { name, nth, status } + * "stage" a turn/exec stage change + */ +function mcodeRun(prompt, { cwd, sessionId, files = [], onProgress, timeoutMs = MCODE_TIMEOUT_MS }) { + return new Promise((resolve) => { + const argv = ["exec", "--cwd", cwd, "--output-format", "stream-json", "--permission", PERMISSION]; + if (sessionId) argv.push("--session", sessionId); + for (const f of files) argv.push("--file", f); + argv.push(prompt); + + const p = spawn(process.execPath, [MCODE_CLI, ...argv], { + shell: false, + windowsHide: true, + stdio: ["ignore", "pipe", "pipe"], + cwd, + }); + + const t0 = Date.now(); + const seen = new Set(); + let toolCount = 0; + let finalOutput = null; + let usage = null; + let streamErr = ""; + let pending = ""; // trailing bytes that do not yet form a line + let settled = false; + let timeoutFired = false; + const timeoutExtra = () => ({ + timedOut: true, + stderr: `still running after ${Math.round(timeoutMs / 1000)}s, force-killed (pid ${p.pid})`, + }); + + const handleLine = (ln) => { + if (!ln.trim()) return; + let ev; + try { ev = JSON.parse(ln); } catch { return; } + const type = ev.type ?? ev.subtype ?? ""; + if (ev.usage) usage = ev.usage; + if (/^(turn|exec)\./.test(type)) onProgress?.("stage", type); + + const tc = ev.item?.toolCall; + if (tc?.name) { + const key = tc.id ?? tc.name; + if (!seen.has(key)) { + seen.add(key); + toolCount++; + onProgress?.("tool", { name: tc.name, nth: toolCount, status: tc.status }); + } + } + const item = ev.item ?? {}; + if (item.type === "text" && (item.content || item.text)) finalOutput = item.content ?? item.text; + if (/^(turn|exec)\.completed$/.test(type)) { + const o = ev.output ?? ev.result?.output ?? ev.item?.content; + if (o) finalOutput = o; + } + }; + + p.stdout.on("data", (d) => { + pending += d.toString("utf8"); + const lines = pending.split("\n"); + pending = lines.pop() ?? ""; + for (const ln of lines) handleLine(ln); + }); + p.stderr.on("data", (d) => (streamErr += d.toString("utf8"))); + + let timer = null; + const settle = (code, extra = {}) => { + if (settled) return; + settled = true; + if (timer) clearTimeout(timer); + if (killGrace) clearTimeout(killGrace); + if (pending.trim()) handleLine(pending); + pending = ""; + resolve({ + output: finalOutput, usage, toolCount, code, timedOut: false, + ...extra, + stderr: extra.stderr ?? streamErr.slice(0, 300), ms: Date.now() - t0, + }); + }; + let killGrace = null; + + // Hard timeout. The worst case in a remote setup is mcode hanging: without + // a timeout the bridge hangs forever and the placeholder message stays on + // "calling xxx…", recoverable only by restarting the process. + // + // Key detail: killTree merely initiates termination, and taskkill is itself + // asynchronous. So do not resolve right away; wait for the close event to + // really arrive (with a 3s backstop) — otherwise the next message races a + // not-yet-dead mcode for the same workspace and session. + timer = setTimeout(() => { + timeoutFired = true; + onProgress?.("timeout", { ms: Date.now() - t0 }); + killTree(p.pid); + killGrace = setTimeout(() => settle(-2, timeoutExtra()), KILL_GRACE_MS); + }, timeoutMs); + + p.on("close", (c) => settle(timeoutFired ? -2 : c, timeoutFired ? timeoutExtra() : {})); + p.on("error", (e) => { streamErr += String(e); settle(-1); }); + }); +} + +// ─────────────────────────── single-instance lock ─────────────────────────── +/** + * A pid lock, guaranteeing that only one bridge runs on a given machine. + * + * Why it is mandatory: two instances would poll the same chat at the same time, + * each reading the same state file and each claiming the same new message — the + * result is mcode running twice, duplicate placeholders, and the two states + * overwriting each other. That is harder to diagnose than a crash, so it is + * blocked at the entry point. + * + * A zombie lock (process force-killed, cleanup never ran) must be reclaimed + * automatically: probe whether the pid is alive, do not rely on file existence. + */ +function isPidAlive(pid) { + if (!Number.isInteger(pid) || pid <= 0) return false; + if (pid === process.pid) return true; + try { + process.kill(pid, 0); // signal 0 only probes existence, it sends nothing + return true; + } catch (e) { + // EPERM means the process exists but the current user may not signal it, so it is alive + return e?.code === "EPERM"; + } +} + +function acquireLock(file = LOCK_FILE) { + if (existsSync(file)) { + const prev = Number(String(readFileSync(file, "utf8")).trim()); + // The lock is ours → let it through idempotently. acquireLock may be called + // from several places, and holding the lock must not make us reject + // ourselves as "some other instance". + if (Number.isInteger(prev) && prev === process.pid) { + return { ok: true, pid: prev, reentrant: true }; + } + if (Number.isInteger(prev) && prev > 0 && isPidAlive(prev)) { + return { ok: false, pid: prev }; + } + // Zombie lock: overwrite directly, do not block startup + try { unlinkSync(file); } catch {} + } + writeFileSync(file, String(process.pid), "utf8"); + return { ok: true, pid: process.pid }; +} + +function releaseLock(file = LOCK_FILE) { + try { + const cur = Number(String(readFileSync(file, "utf8")).trim()); + if (cur === process.pid) unlinkSync(file); // only ever remove our own lock + } catch {} +} + +// ─────────────────────────── state layer ─────────────────────────── +function loadState() { + if (!existsSync(STATE_FILE)) return { chats: {} }; + try { return JSON.parse(readFileSync(STATE_FILE, "utf8")); } catch { return { chats: {} }; } +} + +/** + * Write the state file atomically: stage next to the target, fsync, then rename. + * + * The bridge can be killed at any moment (a user closing the terminal, a reboot, + * the mcode hard timeout). A half-written state file would make loadState() fall + * back to an empty state and the bridge would reprocess the whole conversation, + * re-running mcode on messages it already answered. Rename is atomic on both + * POSIX and Windows, so a reader either sees the old file or the new one. + */ +function saveState(s) { + mkdirSync(dirname(STATE_FILE), { recursive: true }); + const tmp = `${STATE_FILE}.${process.pid}.tmp`; + const fd = openSync(tmp, "w"); + try { + writeFileSync(fd, JSON.stringify(s, null, 2), "utf8"); + try { fsyncSync(fd); } catch {} // best effort; not supported everywhere + } finally { + closeSync(fd); + } + renameSync(tmp, STATE_FILE); +} + +function binding(chatId) { + const st = loadState(); + if (st.chats[chatId]) return { st, b: st.chats[chatId] }; + const cwd = join(WS_ROOT, chatId.replace(/\W/g, "").slice(-14) || "default"); + mkdirSync(cwd, { recursive: true }); + st.chats[chatId] = { cwd, sessionId: null, lastMessageId: null, turns: 0 }; + saveState(st); + return { st, b: st.chats[chatId], fresh: true }; +} + +// ─────────────────────────── message parsing ─────────────────────────── +/** + * Extract text plus attachment file_keys from a Feishu message. + * + * Learned the hard way: Feishu merges "text + image" into msg_type=post, and the + * content lark-cli returns is already markdown (`![Image](img_xxx)`), not + * Feishu's original `{"zh_cn":{"content":[...]}}` structure. So post has to be + * parsed as markdown. + */ +function parseMessage(m) { + const out = { text: "", files: [] }; + const raw = m.content ?? ""; + + if (m.msg_type === "text") { + out.text = raw; + return out; + } + + if (m.msg_type === "image" || m.msg_type === "file") { + let o = {}; + try { o = JSON.parse(raw); } catch {} + const key = o.image_key ?? o.file_key; + if (key) out.files.push({ key, type: m.msg_type === "image" ? "image" : "file" }); + out.text = m.msg_type === "image" ? "[user sent an image]" : "[user sent a file]"; + return out; + } + + if (m.msg_type === "post" || m.msg_type === "rich") { + // Could be markdown, could be Feishu's raw JSON — cover both + let body = raw; + if (raw.trimStart().startsWith("{")) { + try { + const o = JSON.parse(raw); + const lang = o.zh_cn ?? o.en_us ?? Object.values(o)[0] ?? {}; + const nodes = lang.content ?? []; + const parts = []; + for (const n of nodes) { + if (n.tag === "text") parts.push(n.text ?? ""); + else if (n.tag === "img" && n.image_key) { + out.files.push({ key: n.image_key, type: "image" }); + parts.push("[image]"); + } else if (n.tag === "a" && n.href) parts.push(n.text ?? n.href); + else if (n.tag === "at") parts.push(`@${n.user_name ?? "someone"}`); + } + body = parts.join("\n"); + } catch { /* not JSON, so fall through to markdown */ } + } else { + // markdown form: pull out ![Image](key) and [file](key) + const imgRe = /!\[(?:Image|image)?[^\]]*\]\((img_v3_[^)]+)\)/g; + let mt; + while ((mt = imgRe.exec(raw)) !== null) out.files.push({ key: mt[1], type: "image" }); + const fileRe = /\[([^\]]*)\]\((file_v3_[^)]+)\)/g; + while ((mt = fileRe.exec(raw)) !== null) out.files.push({ key: mt[2], type: "file" }); + // strip the markdown image/file syntax, leaving plain text + body = raw + .replace(imgRe, "") + .replace(fileRe, "$1") + .replace(/\[([^\]]*)\]\(https?:\/\/[^)]+\)/g, "$1") + .trim(); + } + if (out.files.length && !body) body = `[user sent ${out.files.length} attachments]`; + out.text = body; + return out; + } + + out.text = `[unsupported message type: ${m.msg_type}]`; + return out; +} + +async function downloadResources(messageId, files, cwd) { + const dir = join(DOWNLOAD_ROOT, messageId.replace(/\W/g, "").slice(-12)); + mkdirSync(dir, { recursive: true }); + const paths = []; + for (const f of files) { + const target = join(dir, f.key); + const r = await lark([ + "im", "+messages-resources-download", + "--message-id", messageId, "--file-key", f.key, + "--type", f.type, "--output", target, "--as", "user", + ]); + if (r.code === 0 && existsSync(target)) paths.push(target); + } + return paths; +} + +// ─────────────────────────── sending layer ─────────────────────────── +/** + * Send a text message. + * + * Must use --content with a JSON body, never --text: + * once --text content contains a newline, going through cmd /c makes the command + * fail silently (the exit code is still 0, but the message is never sent). + * Measured with 46 characters containing 3 newlines: + * --text -> exit=0, output empty, message does not exist ❌ + * --content -> exit=0, all 46 characters delivered ✅ + */ +const textContent = (t) => JSON.stringify({ text: t }); +const msgTypeFlag = ["--msg-type", "text"]; + +const sendMsg = (chatId, text) => + lark(["im", "+messages-send", "--chat-id", chatId, "--as", "bot", + ...msgTypeFlag, "--content", textContent(text)]); + +const replyTo = (messageId, text) => + lark(["im", "+messages-reply", "--message-id", messageId, "--as", "bot", + ...msgTypeFlag, "--content", textContent(text)]); + +/** + * Edit an already-sent text message (placeholder → thinking → final result, + * updated in place). + * + * Note: a bare `api PATCH /open-apis/im/v1/messages/:id` returns + * 230001 "This message is NOT a card", because that endpoint only supports card + * messages. Text messages must use the official `+messages-edit` wrapper, which + * goes through a different endpoint. + */ +const editMsg = (messageId, text) => + lark(["im", "+messages-edit", "--message-id", messageId, "--as", "bot", + ...msgTypeFlag, "--content", textContent(text)]); + +const notify = (chatId, text) => sendMsg(chatId, text); + +/** + * Delivery determination. + * Exit 0 only means the process did not crash — lark-cli can also exit 0 on a + * Feishu-side failure (measured), so when there is JSON, inspect the top-level + * ok, and fall back to the exit code only when there is not. + */ +const sent = (res) => { + if (!res || res.code !== 0) return false; + const t = (res.text || "").trim(); + if (!t.startsWith("{")) return true; + try { return JSON.parse(t)?.ok !== false; } catch { return true; } +}; + +/** + * Edits to the same message must be strictly serial. + * + * Learned the hard way (reproduced, not theoretical): a tool broadcast that + * triggers + * editMsg(phId, hint).catch(() => {}) ← not awaited + * together with the main flow's + * await editMsg(phId, final) + * becomes two concurrent lark-cli processes. Feishu is last-write-wins, so + * whichever lands last wins, and a single lark-cli start takes 1~2 seconds — + * the hint can easily be slower than the final. In 6 reproduction runs, 1 lost + * the final to the hint, leaving the user seeing only "calling read…" with the + * final result never arriving; the demo2.txt run hit exactly this. + * + * The fix: funnel the edits into one promise chain, which is FIFO by nature; + * seal before enqueuing the final, so a late tool event can never rewrite the + * content back to the hint. + */ +function editQueue(editFn = editMsg) { + const failed = (e) => ({ ok: false, code: -1, text: "", err: String(e) }); + let tail = Promise.resolve({ ok: true, text: "", code: 0 }); + let sealed = false; + return { + seal() { sealed = true; }, + /** The final result: still let through after seal, but always queued behind every hint already enqueued */ + push(messageId, text) { + tail = tail.then(() => editFn(messageId, text)).catch(failed); + return tail; + }, + /** Tool broadcast: dropped once the final is enqueued, so a late tool event cannot rewrite the content back to the hint */ + pushHint(messageId, text) { + if (sealed) return; + tail = tail.then(() => editFn(messageId, text)).catch(failed); + }, + /** Wait for the queue to drain, yielding the result of the last edit */ + drain() { return tail; }, + }; +} + +// ─────────────────────────── main flow ─────────────────────────── +/** Quiet tools: read-only/retrieval kinds; announcing them carries no signal and only floods the chat */ +const QUIET_TOOLS = /^(glob|grep|search|list|ls|think|wait|todo)/i; + +/** Delivery retry: after the nth failure wait RETRY_BASE_MS*n, at most MAX_DELIVERY_ATTEMPTS tries */ +const RETRY_BASE_MS = 5000; +const MAX_DELIVERY_ATTEMPTS = 5; +/** How many recent messages one poll fetches. In the normal case a single API call is enough. */ +const PAGE_SIZE = 50; + +/** + * Pick out the messages newer than the watermark from one page of messages + * (oldest first). + * + * This function is pulled out on its own because it once hid a bug that was very + * hard to find: pollOnce used to call `--order asc --page-size 50`, which + * returns the **oldest** 50 messages. Once a conversation grows, new messages + * land on the second page, and if the watermark happens to be the 50th message + * then fresh=[] — the bridge acts as if it saw nothing the user sent, and + * **not a single log line is written**. Worse, as the conversation keeps going + * the watermark gets pushed out of the first page, findIndex=-1, fresh is + * still [], and the bridge is completely deaf. + * + * Returning watermarkFound=false means this page does not cover the watermark; + * the caller must page through to make up the difference, and must never treat + * "not found" as "no new messages". + */ +function selectFresh(pageNewestFirst, lastMessageId) { + const msgs = [...pageNewestFirst].reverse(); // now oldest first + if (!lastMessageId) return { fresh: msgs, watermarkFound: true, total: msgs.length }; + const i = msgs.findIndex((m) => m.message_id === lastMessageId); + return { + fresh: i >= 0 ? msgs.slice(i + 1) : [], + watermarkFound: i >= 0, + total: msgs.length, + }; +} + +/** + * Fetch messages. Normally only the most recent page is pulled (one API call); + * it escalates to a full --page-all walk only when the watermark has been + * pushed off that page (the bridge was stopped for a long time, or someone + * dumped dozens of messages in). Combining --order asc with --page-all also + * works, but pulling the entire history every 2 seconds turns into dozens of + * API calls as soon as there are many messages, which easily hits rate limits. + */ +async function fetchMessages(chatId, lastMessageId) { + const recent = await larkJson([ + "im", "+chat-messages-list", "--chat-id", chatId, + "--order", "desc", "--page-size", String(PAGE_SIZE), "--as", "user", + ]); + if (!recent?.data?.messages?.length) return { messages: [], caughtUp: false, error: !recent }; + + let msgs = recent.data.messages; + let caughtUp = true; + + if (lastMessageId && !msgs.some((m) => m.message_id === lastMessageId)) { + // The watermark is not in the most recent page → we have fallen too far behind and must walk the full history, or messages get lost + console.warn(` ⚠ watermark not in the most recent ${PAGE_SIZE}, escalating to a full paged fetch`); + const full = await larkJson([ + "im", "+chat-messages-list", "--chat-id", chatId, + "--order", "asc", "--page-all", "--as", "user", + ]); + if (full?.data?.messages?.length) { + msgs = full.data.messages; + } else { + caughtUp = false; // the full walk came up empty too; the watermark may have expired + } + } + return { messages: msgs, caughtUp, error: false }; +} + +/** + * Deliver the final result to Feishu: prefer editing the placeholder in place, + * and fall back to a reply when that fails. + * + * Pulled out on its own so a unit test can cover it — this fallback had no test + * coverage at all before, and under negative injection ("delete the fallback") + * the suite still went entirely green, which means it had been running naked. + * The user must never be stuck on a placeholder that stays on "calling xxx…" + * forever. + */ +async function deliverFinal({ q, phId, messageId, finalText, reply = replyTo, log = console.warn }) { + let res; + if (q) { + q.seal(); // cut off late hints + q.push(phId, finalText); // queued behind every hint already enqueued + res = await q.drain(); + } else { + res = await reply(messageId, finalText); + } + if (!sent(res)) { + // Log both stdout and stderr: some lark-cli errors go to stdout, and looking at stderr alone yields a blank page + const detail = [(res?.err || "").trim(), (res?.text || "").trim()] + .filter(Boolean).join(" | ").replace(/\s+/g, " ").slice(0, 200); + log(` ⚠ editing the final result failed code=${res?.code} payload=${finalText.length}ch ` + + `stdout=${(res?.text || "").length}B stderr=${(res?.err || "").length}B` + + `${detail ? " :: " + detail : " :: (no output)"}`); + if (process.env.MCODE_FEISHU_BRIDGE_DEBUG) { + writeFileSync(join(TMP, "bridge-final-debug.json"), + JSON.stringify({ phId, messageId, finalText }, null, 2), "utf8"); + log(` 🔍 payload written to ${join(TMP, "bridge-final-debug.json")}`); + } + res = await reply(messageId, finalText); + if (!sent(res)) { + log(` ⚠ reply failed too code=${res?.code} payload=${finalText.length}ch ` + + `stdout=${(res?.text || "").length}B stderr=${(res?.err || "").length}B ` + + `:: ${[(res?.err || "").trim(), (res?.text || "").trim()].filter(Boolean).join(" | ").slice(0, 200)}`); + } + } + return { res, ok: sent(res) }; +} + +/** + * Classify what went wrong while fetching messages, **classify only, do not + * print**. + * + * Why classify instead of throwing straight away: that way "whether anything + * went wrong, and which kind" can be held by a unit test (deleting the warning + * used to leave all 80 tests green anyway). What actually makes it loud is the + * throw in collectFresh below. + */ +function reportFetchProblem({ error, watermarkFound, watermark }) { + if (error) { + return { + ok: false, reason: "fetch-error", + message: "Failed to fetch messages: lark-cli returned no usable data. Check whether the token is expired, plus network and permissions.", + }; + } + if (!watermarkFound) { + return { + ok: false, reason: "no-watermark", + message: `The watermark ${watermark} is not in this page, so this round is skipped to avoid duplicates or lost messages.` + + " (usually the bridge was stopped long enough that the history got pushed out of the most recent page; check logs/bridge.log)", + }; + } + return { ok: true, reason: "ok", message: "" }; +} + +class BridgeFetchError extends Error { + constructor(problem) { + super(`[${problem.reason}] ${problem.message}`); + this.name = "BridgeFetchError"; + this.reason = problem.reason; + } +} + +/** + * Collect the messages newer than the watermark. + * + * **If it cannot be fetched, throw — never silently return an empty array.** + * A throw is used rather than a status code that has to be remembered and + * checked, because "forgot to check the return value" is a failure mode tests + * cannot catch (delete the single line `if (problem !== "ok") return 0` under + * negative injection and the whole suite still goes green). Once thrown, the + * only exit is the catch in the watch loop, which already logs — being loud + * is guaranteed structurally, not by discipline. + */ +async function collectFresh(chatId, lastMessageId, { fetch = fetchMessages } = {}) { + const { messages = [], caughtUp, error } = await fetch(chatId, lastMessageId); + + // Treat every empty page as "failed to fetch messages", without depending on + // the caller remembering to set the error flag. Depending on another function + // to set a flag and depending on the caller to check a return value are the + // same species of fragility. + const problem = reportFetchProblem({ + error: !!error || messages.length === 0, + watermarkFound: messages.length > 0 && selectFresh(messages, lastMessageId).watermarkFound, + watermark: lastMessageId, + }); + if (!problem.ok) throw new BridgeFetchError(problem); + return { fresh: messages.length ? selectFresh(messages, lastMessageId).fresh : [], caughtUp }; +} + +async function pollOnce(chatId) { + const { st, b } = binding(chatId); + + // Flush the previous round's undelivered result before considering new messages + await flushPending(chatId); + + const { fresh } = await collectFresh(chatId, b.lastMessageId); + + if (fresh.length > 50) console.warn(` ⚠ ${fresh.length} new messages backed up at once`); + + let handled = 0; + for (const m of fresh) { + if (m.deleted) { b.lastMessageId = m.message_id; continue; } + const who = m.sender?.sender_type; + if (who !== "user") continue; // ignore our own messages + if (m.message_id === b.lastMessageId) continue; + + const { text, files } = parseMessage(m); + if (!text.trim() && !files.length) { + // Unknown type: warn but do not advance the watermark, so it can still be replayed once the parser catches up + console.warn(` ⚠ skipping an unparsable message type=${m.msg_type} id=${m.message_id}`); + continue; + } + + // ① Reply with a placeholder immediately, keeping the "received → some reaction" wait as short as possible + const ph = await larkJson([ + "im", "+messages-send", "--chat-id", chatId, "--as", "bot", + ...msgTypeFlag, "--content", textContent("🧠 Got it, thinking…"), + ]); + const phId = ph?.data?.message_id ?? ph?.data?.message?.message_id ?? null; + if (!phId) console.warn(" ⚠ placeholder message failed to send, falling back to reply"); + + // From the moment the placeholder goes out, every write to this message goes through one serial queue + const q = phId ? editQueue() : null; + + // ② Download the attachments + const paths = files.length ? await downloadResources(m.message_id, files, b.cwd) : []; + if (paths.length && q) { + q.pushHint(phId, `📎 Received ${paths.length} attachments, starting…`); + } + + // ③ Run mcode. The placeholder is updated in place, no chat spam. + // Every edit goes through q (the serial queue); that is the key to fixing + // "the hint overwrites the final result". + let toolCount = 0; + let lastHint = ""; + const started = Date.now(); + const r = await mcodeRun(text, { + cwd: b.cwd, sessionId: b.sessionId, files: paths, + timeoutMs: MCODE_TIMEOUT_MS, + onProgress: (kind, detail) => { + if (kind === "timeout" && q) { + q.pushHint(phId, `⏱ Over ${Math.round(MCODE_TIMEOUT_MS / 60000)} minutes, force-terminating…`); + return; + } + if (kind !== "tool" || !q) return; + toolCount = detail.nth; + if (QUIET_TOOLS.test(detail.name)) return; + const hint = `⚙️ Calling \`${detail.name}\`…`; + if (hint === lastHint) return; + lastHint = hint; + q.pushHint(phId, hint); // enqueueing is enough, do not await; the queue guarantees ordering + }, + }); + + // ④ The final result + const secs = ((Date.now() - started) / 1000).toFixed(1); + const bits = [`⏱ ${secs}s`]; + if (r.usage?.total_tokens) bits.push(`${r.usage.total_tokens} tokens`); + if (toolCount) bits.push(`${toolCount} tool calls`); + + const finalText = r.timedOut + ? `⏱ **timed out and was force-terminated** (${secs}s)\n\n` + + `This mcode turn ran for ${secs}s without finishing, and the process has been killed.\n` + + `Common causes: a tool stuck on an interactive prompt, a command waiting for input, or a hung network request.\n\n` + + `${toolCount} tool calls had already completed, and their side effects on the workspace are still there,` + + `so you can ask me to check the current state, or rephrase and try again.\n\n---\n${bits.join(" · ")}` + : r.output + ? `${r.output}\n\n---\n${bits.join(" · ")}` + : `❌ Processing failed (${secs}s)\n\`\`\`\n${(r.stderr || "").slice(0, 300)}\n\`\`\``; + + // ④ Delivery: prefer editing the placeholder in place, automatically fall + // back to a reply if the edit fails + const { ok } = await deliverFinal({ q, phId, messageId: m.message_id, finalText }); + + // ⑤ Record. The watermark advances whether or not delivery succeeded, and an + // undelivered result is stored separately in b.pending. + // See the recordOutcome comment — this used to be the source of a + // runaway loop that flooded the chat. + recordOutcome(b, { messageId: m.message_id, ok, phId, finalText, paths }); + saveState(st); + handled++; + console.log(` ${ok ? "✓" : "⏳"} ${m.create_time} ${text.slice(0, 40)} → ${secs}s` + + (ok ? "" : "(delivery failed, queued; only delivery is retried, mcode is not re-run)")); + } + return handled; +} + +/** Linear backoff: after the nth failure wait base*n milliseconds */ +const backoffMs = (attempts, base = RETRY_BASE_MS) => base * Math.max(1, attempts); +/** Give up once the consecutive failures hit the limit; never retry forever */ +const shouldGiveUp = (attempts, max = MAX_DELIVERY_ATTEMPTS) => attempts >= max; + +/** + * Record the outcome of handling one message. + * + * The most critical contract: **the watermark advances whether or not delivery + * succeeded**. It used to be written as "do not advance the watermark when + * delivery fails", intending to re-deliver on the next round — but as soon as + * delivery kept failing, the whole pollOnce re-ran from the start, re-running + * mcode and re-sending a placeholder every round, flooding Feishu. A failed + * result goes into b.pending instead, and only delivery is retried, never a + * re-run of mcode. + * + * Extracted into a pure function so a unit test can hold this contract. + */ +function recordOutcome(b, { messageId, ok, phId = null, finalText, paths = [], now = Date.now() }) { + b.lastMessageId = messageId; + b.turns = (b.turns ?? 0) + 1; + if (paths.length) b.lastFiles = paths; + b.pending = ok ? null : { messageId, phId, finalText, attempts: 1, nextTryAt: now + backoffMs(1) }; + return b; +} + +/** + * Re-deliver the previous round's undelivered result. + * Only delivery is retried, mcode is not re-run, and no new placeholder is sent — + * that placeholder stuck on "calling xxx…" is simply edited in place into the + * final result. + */ +async function flushPending(chatId) { + const { st, b } = binding(chatId); + const p = b.pending; + if (!p) return 0; + if (Date.now() < (p.nextTryAt ?? 0)) return 0; + + const q = p.phId ? editQueue() : null; + const { ok } = await deliverFinal({ q, phId: p.phId, messageId: p.messageId, finalText: p.finalText }); + if (ok) { + b.pending = null; + saveState(st); + console.log(` ✅ re-delivery succeeded (attempt ${p.attempts})`); + return 1; + } + + p.attempts = (p.attempts ?? 1) + 1; + if (shouldGiveUp(p.attempts)) { + b.pending = null; + await notify(chatId, `⚠️ The result could not be delivered in ${p.attempts} consecutive attempts, giving up. The original result:\n\n${p.finalText}`); + saveState(st); + console.error(` ✗ delivery failed ${p.attempts} times in a row, giving up (never retried forever)`); + return 0; + } + p.nextTryAt = Date.now() + backoffMs(p.attempts); + b.pending = p; + saveState(st); + console.warn(` ⏳ delivery retry ${p.attempts}/${MAX_DELIVERY_ATTEMPTS}, in ${Math.round((p.nextTryAt - Date.now()) / 1000)}s`); + return 0; +} + +// ─────────────────────────── entry point ─────────────────────────── +async function main() { + // Only watch mode needs exclusivity: once mode is one-shot, and the lock would only get in the way + if (WATCH) { + const lock = acquireLock(); + if (!lock.ok) { + console.error(`✗ A bridge is already running (pid ${lock.pid}).`); + console.error(` Two instances would claim the same message and overwrite each other's state; stop the old one first:`); + console.error(` node mcode-feishu-bridge.mjs --stop`); + process.exit(3); + } + const bye = () => { releaseLock(); }; + process.on("exit", bye); + process.on("SIGINT", () => { console.log("\nExiting…"); releaseLock(); process.exit(0); }); + process.on("SIGTERM", () => { releaseLock(); process.exit(0); }); + process.on("uncaughtException", (e) => { + console.error(`Uncaught exception: ${e?.stack || e}`); + releaseLock(); + process.exit(1); + }); + } + + mkdirSync(WS_ROOT, { recursive: true }); + mkdirSync(DOWNLOAD_ROOT, { recursive: true }); + + const { b } = binding(CHAT_ID); + console.log(`chat ${CHAT_ID}`); + console.log(`workspace ${b.cwd}`); + console.log(`session ${b.sessionId ?? "(created on the first turn)"}`); + console.log(`mode ${WATCH ? `watch ${INTERVAL}ms` : "once"}`); + console.log(`timeout ${Math.round(MCODE_TIMEOUT_MS / 1000)}s`); + console.log(`lock ${WATCH ? `pid ${process.pid}` : "(no lock in once mode)"}\n`); + + if (!WATCH) { + const n = await pollOnce(CHAT_ID); + console.log(n ? `\nHandled ${n}` : "No new messages"); + return; + } + + let running = true; + process.on("SIGINT", () => { running = false; }); + + console.log(`Watching, Ctrl+C to exit\n`); + while (running) { + try { + await pollOnce(CHAT_ID); + } catch (e) { + console.error(`Polling error: ${e.message}`); + } + await new Promise((r) => setTimeout(r, INTERVAL)); + } + releaseLock(); + console.log("Exited."); +} + +/** Stop the running watch instance (kills the process tree using the pid in the lock) */ +function stopRunning() { + if (!existsSync(LOCK_FILE)) { + console.log("No running instance (no lock file)."); + return; + } + const pid = Number(String(readFileSync(LOCK_FILE, "utf8")).trim()); + if (!isPidAlive(pid)) { + console.log(`The lock is a zombie (pid ${pid} no longer exists); cleaning it up.`); + releaseLock(); + return; + } + killTree(pid); + try { unlinkSync(LOCK_FILE); } catch {} + console.log(`Stopped pid ${pid}.`); +} + +// With MCODE_FEISHU_BRIDGE_TEST=1 this only exports and never starts watching, so +// the unit tests can import the real implementation instead of testing a copy of +// the logic (a copy is always green, which is the same as not testing at all). +export { editQueue, sent, parseMessage, textContent, QUIET_TOOLS, deliverFinal, + backoffMs, shouldGiveUp, recordOutcome, killTree, isPidAlive, + acquireLock, releaseLock, mcodeRun, selectFresh, reportFetchProblem, collectFresh, + resolveMcodeCli, resolveLarkExe, DATA_DIR, IS_WIN, + MCODE_TIMEOUT_MS, MAX_DELIVERY_ATTEMPTS, RETRY_BASE_MS, PAGE_SIZE }; + +if (process.env.MCODE_FEISHU_BRIDGE_TEST === "1") { + // Unit-test mode: export only, do not start +} else if (args.includes("--stop")) { + stopRunning(); +} else { + main(); +} diff --git a/plugins/antianqi/mcode-feishu-bridge/scripts/mcode-feishu-bridge.test.mjs b/plugins/antianqi/mcode-feishu-bridge/scripts/mcode-feishu-bridge.test.mjs new file mode 100644 index 00000000..43a19560 --- /dev/null +++ b/plugins/antianqi/mcode-feishu-bridge/scripts/mcode-feishu-bridge.test.mjs @@ -0,0 +1,447 @@ +/** + * editQueue 单测 —— 直接 import 真实实现,不是抄一份。 + * + * 负向对照设计(对应实测 bug): + * 用例 A 复刻修复前的并发写法(hint 不 await + final await), + * 断言"最终结果必须是消息里最后落盘的内容" → 必然失败 + * 用例 B 用真实 editQueue 做同样的事 → 必须通过 + * 如果哪天有人把 editQueue 改回并发,这个 A/B 对照会立刻翻脸。 + */ +process.env.MCODE_FEISHU_BRIDGE_TEST = "1"; +import { readFileSync, writeFileSync, existsSync, unlinkSync, mkdtempSync, rmSync } from "node:fs"; +import { join, dirname } from "node:path"; +import { tmpdir } from "node:os"; +import { spawn } from "node:child_process"; +import { fileURLToPath } from "node:url"; +const { editQueue, sent, deliverFinal, backoffMs, shouldGiveUp, recordOutcome, + acquireLock, releaseLock, isPidAlive, killTree, mcodeRun, selectFresh, resolveMcodeCli, + reportFetchProblem, collectFresh, PAGE_SIZE, + MCODE_TIMEOUT_MS, MAX_DELIVERY_ATTEMPTS, RETRY_BASE_MS } = + await import("./mcode-feishu-bridge.mjs"); + +let pass = 0, fail = 0, skipped = 0; +const results = []; +function check(name, ok, detail = "") { + if (ok) { pass++; results.push(` ok ${name}`); } + else { fail++; results.push(` FAIL ${name}${detail ? " <- " + detail : ""}`); } +} +/** 宿主缺少前置条件时跳过,而不是失败——CI 上本来就没有 mcode。 */ +function skip(name) { skipped++; results.push(` skip ${name}`); } + +/** + * 假飞书:内容 = 最后落盘的那次写入。 + * @param durations 每次写入的耗时(ms),用来制造乱序 + */ +function fakeFeishu(durations) { + let n = 0; + const state = { content: null, order: [] }; + const timers = []; + const edit = (messageId, text) => new Promise((res) => { + const d = durations[n++] ?? 10; + timers.push(setTimeout(() => { + state.content = text; // 后写覆盖 + state.order.push(text); + res({ code: 0, text: JSON.stringify({ ok: true }), ok: true }); + }, d)); + }); + return { edit, state, clear: () => timers.forEach(clearTimeout) }; +} + +const HINT = "⚙️ 正在调用 `read`…"; +const FINAL = "✅ demo2.txt = final test\n\n---\n⏱ 12.3s · 1335 tokens · 2 次工具调用"; + +// ── A. 负向对照:修复前的并发写法,hint 比 final 慢 → final 必然被盖掉 ── +{ + const { edit, state, clear } = fakeFeishu([80, 10]); // hint 80ms,final 10ms + const flying = edit("m1", HINT).catch(() => {}); // 不 await(= 旧代码) + await edit("m1", FINAL); // await(= 旧代码) + await flying; + clear(); + check("A 旧并发写法:final 确实会被慢 hint 盖掉(负向对照成立)", + state.content === HINT, + `期望 content===HINT,实际 content=${JSON.stringify(state.content)}`); +} + +// ── B. 真实 editQueue:同样"慢 hint + 快 final",final 必须赢 ── +{ + const { edit, state, clear } = fakeFeishu([80, 10]); + const q = editQueue(edit); + q.pushHint("m1", HINT); + q.seal(); + q.push("m1", FINAL); + await q.drain(); + clear(); + check("B 串行队列:慢 hint 也排在 final 之前,final 赢", + state.content === FINAL, + `实际 content=${JSON.stringify(state.content)}`); + check("B 顺序:hint 先落盘、final 后落盘", + JSON.stringify(state.order) === JSON.stringify([HINT, FINAL]), + `order=${JSON.stringify(state.order)}`); +} + +// ── C. seal 之后到达的 hint 必须被丢弃 ── +{ + let calls = 0; + const edit = async (_id, text) => { calls++; return { code: 0, text: JSON.stringify({ ok: true }), content: text }; }; + const q = editQueue(edit); + q.push("m1", FINAL); + q.seal(); + q.pushHint("m1", HINT); // 晚到的 tool 事件 + await q.drain(); + check("C seal 后 hint 被丢弃(只发 1 次)", calls === 1, `calls=${calls}`); +} + +// ── D. seal 之前入队的 hint 必须放行(不能把中间状态也吞掉) ── +{ + let calls = 0; + const edit = async () => { calls++; return { code: 0, text: JSON.stringify({ ok: true }) }; }; + const q = editQueue(edit); + q.pushHint("m1", HINT); + q.seal(); + q.push("m1", FINAL); + await q.drain(); + check("D seal 前入队的 hint 正常执行(共 2 次)", calls === 2, `calls=${calls}`); +} + +// ── E. 中途一次编辑抛错,不能卡死队列(失败的在中间,final 仍要落盘) ── +{ + const state = { content: null }; + const timers = []; + let n = 0; + const edit = (id, text) => { + const i = n++; + if (i === 1) return Promise.reject(new Error("boom")); // 中间那次失败 + return new Promise((res) => timers.push(setTimeout(() => { + state.content = text; + res({ code: 0, text: "{}" }); + }, 5))); + }; + const q = editQueue(edit); + q.pushHint("m1", HINT); + q.pushHint("m1", "⚙️ 正在调用 `write`…"); // 这次会抛错 + q.seal(); + q.push("m1", FINAL); + await q.drain(); + timers.forEach(clearTimeout); + check("E 中途抛错不卡死队列,final 仍落盘", state.content === FINAL, `content=${JSON.stringify(state.content)}`); +} + +// ── F. sent() 送达判定 ── +check("F exit!=0 → 不算送达", sent({ code: 1, text: "" }) === false); +check("F exit0 + JSON ok:false → 不算送达", sent({ code: 0, text: '{"ok":false}' }) === false); +check("F exit0 + JSON ok:true → 算送达", sent({ code: 0, text: '{"ok":true}' }) === true); +check("F exit0 + 无 JSON → 算送达", sent({ code: 0, text: "plain text" }) === true); +check("F undefined → 不算送达", sent(undefined) === false); + +// ── E2. final 自己失败时,drain 返回失败标记(不 throw),让调用方能退回 reply ── +{ + const q = editQueue(() => Promise.reject(new Error("edit failed"))); + q.seal(); + q.push("m1", FINAL); + const r = await q.drain(); + check("E2 final 失败 → drain 返回失败标记而非 throw", + r && r.ok === false && r.code === -1, `r=${JSON.stringify(r)}`); + check("E2 失败标记会被 sent() 判为未送达", sent(r) === false); +} + +// ── G. deliverFinal:edit 成功就不该多发一条 reply ── +{ + let edits = 0, replies = 0; + const q = editQueue(async () => { edits++; return { code: 0, text: '{"ok":true}' }; }); + const reply = async () => { replies++; return { code: 0, text: '{"ok":true}' }; }; + const r = await deliverFinal({ q, phId: "m1", messageId: "u1", finalText: FINAL, reply, log: () => {} }); + check("G edit 成功 → 不发 reply", edits === 1 && replies === 0, `edits=${edits} replies=${replies}`); + check("G edit 成功 → ok=true", r.ok === true); +} + +// ── H. edit 失败 → 必须退回 reply,用户不能卡在占位上 ── +{ + let replies = 0, replied = null; + const q = editQueue(async () => ({ code: 1, text: "", err: "edit blew up" })); + const reply = async (mid, text) => { replies++; replied = { mid, text }; return { code: 0, text: '{"ok":true}' }; }; + const r = await deliverFinal({ q, phId: "m1", messageId: "u1", finalText: FINAL, reply, log: () => {} }); + check("H edit 失败 → 退回 reply", replies === 1, `replies=${replies}`); + check("H 退回的 reply 指向原用户消息且内容完整", + replied && replied.mid === "u1" && replied.text === FINAL, `replied=${JSON.stringify(replied)}`); + check("H 兜底成功 → ok=true", r.ok === true); +} + +// ── I. edit 与 reply 都失败 → 必须报未送达,让水位不推进 ── +{ + const q = editQueue(async () => ({ code: 1, text: "", err: "edit blew up" })); + const reply = async () => ({ code: 0, text: '{"ok":false}' }); + const r = await deliverFinal({ q, phId: "m1", messageId: "u1", finalText: FINAL, reply, log: () => {} }); + check("I 两处都失败 → ok=false(水位不推进,下轮重试)", r.ok === false); +} + +// ── J. 没有占位消息(发送占位就失败)→ 直接 reply,不做无谓的 edit ── +{ + let edits = 0, replies = 0; + const reply = async () => { replies++; return { code: 0, text: '{"ok":true}' }; }; + const r = await deliverFinal({ q: null, phId: null, messageId: "u1", finalText: FINAL, reply, log: () => {} }); + check("J 无占位 → 直接 reply 一次", replies === 1 && edits === 0, `replies=${replies} edits=${edits}`); + check("J 无占位 → ok=true", r.ok === true); +} + +// ── K. deliverFinal 必须把队列封住:它返回后再来的 hint 不能改写最终结果 ── +{ + const written = []; + const q = editQueue(async (_id, text) => { written.push(text); return { code: 0, text: '{"ok":true}' }; }); + await deliverFinal({ q, phId: "m1", messageId: "u1", finalText: FINAL, reply: async () => ({ code: 0, text: "{}" }), log: () => {} }); + q.pushHint("m1", HINT); // mcode 已收工后才到的 tool 事件 + await q.drain(); + check("K deliverFinal 后队列已封,晚到 hint 不改写结果", + written.length === 1 && written[0] === FINAL, + `written=${JSON.stringify(written)}`); +} + +// ── L. 退避必须是单调递增的,不能是 0(0 = 立刻重试 = 刷屏) ── +check("L 退避随次数递增", backoffMs(1) < backoffMs(2) && backoffMs(2) < backoffMs(3), + `b1=${backoffMs(1)} b2=${backoffMs(2)} b3=${backoffMs(3)}`); +check("L 退避永不为 0(否则等于无限立即重试)", + [1, 2, 3, 10, 999].every(n => backoffMs(n) > 0)); +check("L 退避下限是 RETRY_BASE_MS", backoffMs(1) === RETRY_BASE_MS, `b1=${backoffMs(1)}`); +check("L attempts=0 也不会退化成 0", backoffMs(0) === RETRY_BASE_MS, `b0=${backoffMs(0)}`); + +// ── M. 必须有放弃上限:这是"无限重试刷屏"的唯一防线 ── +check("M 上限存在且有限", Number.isFinite(MAX_DELIVERY_ATTEMPTS) && MAX_DELIVERY_ATTEMPTS > 0 && MAX_DELIVERY_ATTEMPTS <= 20, + `max=${MAX_DELIVERY_ATTEMPTS}`); +check("M 到上限必须放弃", shouldGiveUp(MAX_DELIVERY_ATTEMPTS) === true); +check("M 超过上限也放弃", shouldGiveUp(MAX_DELIVERY_ATTEMPTS + 1) === true && shouldGiveUp(999) === true); +check("M 上限之前不能放弃", + ![1, 2, MAX_DELIVERY_ATTEMPTS - 1].some(n => shouldGiveUp(n)), + `max=${MAX_DELIVERY_ATTEMPTS}`); + +// ── N. recordOutcome:水位必须无条件推进,否则会刷屏死循环 ── +{ + const b1 = { lastMessageId: "old", turns: 3, pending: null }; + recordOutcome(b1, { messageId: "u1", ok: true, phId: "p1", finalText: "F" }); + check("N 送达成功 → 水位推进、pending 清空", + b1.lastMessageId === "u1" && b1.pending === null, JSON.stringify(b1)); +} +{ + const b2 = { lastMessageId: "old", turns: 3, pending: null }; + recordOutcome(b2, { messageId: "u1", ok: false, phId: "p1", finalText: "F", now: 1000 }); + check("N 送达失败 → 水位照样推进(否则整条 pollOnce 重跑 = 刷屏)", + b2.lastMessageId === "u1", `lastMessageId=${b2.lastMessageId}`); + check("N 送达失败 → 结果挂进 pending 等待补投", + b2.pending && b2.pending.messageId === "u1" && b2.pending.finalText === "F" && b2.pending.phId === "p1", + JSON.stringify(b2.pending)); + check("N 送达失败 → pending 带退避时间,不是立刻重试", + b2.pending.nextTryAt === 1000 + RETRY_BASE_MS, `nextTryAt=${b2.pending.nextTryAt}`); + check("N 送达失败 → pending 记录了已试次数", b2.pending.attempts === 1); +} +{ + // 曾经有 pending 时又被一次成功投递清掉,不能留脏状态 + const b3 = { lastMessageId: "old", turns: 0, pending: { messageId: "u0", attempts: 3 } }; + recordOutcome(b3, { messageId: "u2", ok: true, finalText: "F" }); + check("N 成功时清掉旧 pending", b3.pending === null); +} +{ + // 附件路径要被记下来(下载目录要能追溯) + const b4 = { lastMessageId: "old", turns: 0 }; + recordOutcome(b4, { messageId: "u1", ok: true, paths: ["a.jpg", "b.pdf"] }); + check("N 附件路径被记录", Array.isArray(b4.lastFiles) && b4.lastFiles.length === 2); + check("N 轮次自增", b4.turns === 1, `turns=${b4.turns}`); +} + +// ── O. 静态护栏:源码里绝不能再出现 cmd.exe / shell:true ── +{ + const src = readFileSync(new URL("./mcode-feishu-bridge.mjs", import.meta.url), "utf8"); + // 去掉块注释和行注释再查,否则文档里提到 cmd.exe 会误报 + const codeOnly = src.replace(/\/\*[\s\S]*?\*\//g, "").replace(/^\s*\/\/.*$/gm, ""); + check("O 源码里没有 cmd.exe(回到 shell 转发 = 双重解析 bug 复现)", + !/cmd\.exe/.test(codeOnly), "发现 cmd.exe"); + check("O 源码里没有 shell: true", !/shell\s*:\s*true/.test(codeOnly), "发现 shell:true"); + // 逐个 spawn 调用的参数里都必须没有 shell:true + const spawnBlocks = codeOnly.split(/spawn\(/).slice(1); + const anyShellTrue = spawnBlocks.some(b => /shell\s*:\s*true/.test(b.slice(0, 200))); + check("O 每个 spawn 的参数里都没有 shell:true", !anyShellTrue); + check("O 确实存在 spawn 调用(否则上面的检查是空转)", spawnBlocks.length >= 2, `spawn 次数=${spawnBlocks.length}`); +} + +// ── P. 单实例锁:活着就拒绝,僵尸锁要能回收,只删自己的锁 ── +{ + const f = join(tmpdir(), `fmb-lock-${process.pid}-a.lock`); + const l1 = acquireLock(f); + check("P 首次加锁成功", l1.ok === true && l1.pid === process.pid); + + const l2 = acquireLock(f); + check("P 同一进程重复加锁视为自己(幂等,不阻塞)", l2.ok === true, JSON.stringify(l2)); + + // 用一个真起着的子进程当"别人的锁"。 + // 曾经写死 pid 1 一定活着——Windows 上根本没有 PID 1,会被当成僵尸锁回收, + // 测试就假绿了。所以这里必须真起一个进程。 + const child = spawn(process.execPath, ["-e", "setTimeout(()=>{},60000)"], + { shell: false, windowsHide: true, stdio: "ignore" }); + releaseLock(f); + await new Promise(r => setTimeout(r, 300)); // 等它真的起来 + check("P 前置:子进程确实活着", isPidAlive(child.pid) === true, `pid=${child.pid}`); + + writeFileSync(f, String(child.pid)); + const l3 = acquireLock(f); + check("P 别的活进程持锁 → 拒绝启动", l3.ok === false && l3.pid === child.pid, JSON.stringify(l3)); + + // 僵尸锁:pid 是个几乎不可能存在的大数 + writeFileSync(f, "4194300"); + const l4 = acquireLock(f); + check("P 僵尸锁(进程已死)→ 自动回收并接管", l4.ok === true, JSON.stringify(l4)); + + // 只删自己的锁 + writeFileSync(f, "999999"); + releaseLock(f); + check("P releaseLock 不误删别人的锁", existsSync(f), "锁被误删"); + + child.kill(); + await new Promise(r => setTimeout(r, 200)); + check("P 清理:子进程已回收", isPidAlive(child.pid) === false); + try { unlinkSync(f); } catch {} + check("P 锁文件已清理", existsSync(f) === false); + + check("P isPidAlive 对自己为真", isPidAlive(process.pid) === true); + check("P isPidAlive 对明显不存在的 pid 为假", isPidAlive(2147483647) === false); + check("P isPidAlive 对垃圾输入为假", + [0, -1, NaN, "abc", null, undefined, 1.5].every(v => isPidAlive(v) === false)); +} + +// ── Q. mcode hard timeout: one real run with a 1.5s budget (a real turn needs ~11s). +// Skipped when this host has no mcode, which is the normal case on CI. ── +{ + check("Q 默认超时有值且合理(1~60 分钟)", + MCODE_TIMEOUT_MS >= 60000 && MCODE_TIMEOUT_MS <= 3600000, `default=${MCODE_TIMEOUT_MS}`); + + if (!resolveMcodeCli()) { + skip("Q mcode 未安装,跳过真实超时用例(CI 上没有 mcode)"); + } else { + const wd = mkdtempSync(join(tmpdir(), "fmb-timeout-")); + let sawTimeoutEvent = false; + const t0 = Date.now(); + const r = await mcodeRun("写一句话:你好", { cwd: wd, timeoutMs: 1500, onProgress: (k) => { if (k === "timeout") sawTimeoutEvent = true; } }); + const elapsed = Date.now() - t0; + check("Q 超时被标记", r.timedOut === true, JSON.stringify({ timedOut: r.timedOut, code: r.code })); + check("Q 退出码是 -2(专用超时码,不是普通失败)", r.code === -2, `code=${r.code}`); + check("Q onProgress 收到了 timeout 事件", sawTimeoutEvent === true); + check("Q 超时提示含秒数,便于用户判断", /\d+s/.test(r.stderr || ""), JSON.stringify(r.stderr)); + check("Q 及时返回(不等到 mcode 自己结束)", elapsed < 8000, `elapsed=${elapsed}ms`); + // 清理失败也要能跑完(被杀进程的句柄可能还没释放),但要能看出是哪种情况 + let removed = true; + try { rmSync(wd, { recursive: true, force: true, maxRetries: 5, retryDelay: 200 }); } + catch { removed = false; } + check("Q 超时后子进程已退出,临时目录可删(没留僵尸句柄)", removed === true, + "目录删不掉,说明 mcode 进程还占着句柄"); + } +} + +// ── T. selectFresh:水位切分新消息 ── +{ + const mk = (n) => ({ message_id: `om_${n}`, create_time: `t${n}` }); + const pageDesc = [mk(10), mk(9), mk(8), mk(7)]; // desc(最新在前) + const r1 = selectFresh(pageDesc, "om_8"); + check("T desc 页被翻转成正序", + r1.fresh.map(m => m.message_id).join(",") === "om_9,om_10", + r1.fresh.map(m => m.message_id).join(",")); + check("T 命中水位时 watermarkFound=true", r1.watermarkFound === true); + + const r2 = selectFresh(pageDesc, null); + check("T 没有水位(首轮)→ 全部当新消息", r2.fresh.length === 4 && r2.watermarkFound === true); + + const r3 = selectFresh(pageDesc, "om_10"); // 水位就是最新那条 + check("T 水位=最新 → 没有新消息", r3.fresh.length === 0 && r3.watermarkFound === true); + + const r4 = selectFresh(pageDesc, "om_1"); // 水位太老,不在这一页 + check("T 水位不在页内 → fresh 为空", r4.fresh.length === 0); + check("T 水位不在页内 → watermarkFound=false(必须触发翻页)", r4.watermarkFound === false, + "这里如果错成 true,bridge 就会静默丢消息"); +} + +// ── U. 回归:真实场景复现(会话消息超过一页,bridge 会变聋) ── +{ + // 复刻这次的现场:--order asc --page-size 50 拿到的是最老 50 条, + // 水位恰好是最后一条 → fresh=[],用户 16:41 的 hello 被静默吞掉。 + const oldest = Array.from({ length: 50 }, (_, i) => ({ + message_id: `om_old${i}`, create_time: `2026-10-01 14:${String(45 + i).padStart(2, "0")}`, + })); + // desc 页 = 最新在前。bridge 发的占位比用户消息晚,所以 placeholder 更新。 + const realNewest = [ + { message_id: "om_placeholder", create_time: "2026-10-01 16:41" }, + { message_id: "om_hello", create_time: "2026-10-01 16:41" }, + ]; + const watermark = "om_old49"; + + // 旧行为:直接对 asc 页做 findIndex + const oldIdx = oldest.findIndex(m => m.message_id === watermark); + const oldFresh = oldIdx >= 0 ? oldest.slice(oldIdx + 1) : []; + check("U 旧行为确实漏掉了(负向对照成立)", + oldFresh.length === 0 && realNewest.length > 0, `oldFresh=${oldFresh.length}`); + + // 新行为:先取最近一页 desc,命中不了才升级全量 + const page1 = selectFresh(realNewest, watermark); + check("U 最近一页没命中水位 → watermarkFound=false(触发升级)", + page1.watermarkFound === false); + const escalated = selectFresh([...realNewest, ...[...oldest].reverse()], watermark); + check("U 升级全量后能取到那两条新消息", + escalated.fresh.map(m => m.message_id).join(",") === "om_hello,om_placeholder", + escalated.fresh.map(m => m.message_id).join(",")); +} + +// ── V. 源码静态护栏:取消息必须走 desc + 水位升级,不能退回 asc 单页 ── +{ + const src = readFileSync(new URL("./mcode-feishu-bridge.mjs", import.meta.url), "utf8"); + const codeOnly = src.replace(/\/\*[\s\S]*?\*\//g, "").replace(/^\s*\/\/.*$/gm, ""); + const listCalls = [...codeOnly.matchAll(/\+chat-messages-list[\s\S]{0,320}?\]/g)].map(m => m[0]); + check("V 存在取消息调用", listCalls.length >= 1, `找到 ${listCalls.length} 处`); + const bareAsc = listCalls.filter(c => /--order",\s*"asc"/.test(c) && !/--page-all/.test(c)); + check("V 没有「asc 且不带 --page-all」的取消息调用(那正是本次失聪的根因)", + bareAsc.length === 0, `${bareAsc.length} 处`); + const hasDesc = listCalls.some(c => /--order",\s*"desc"/.test(c) && /--page-size/.test(c)); + check("V 常态路径是 desc + 限页(一次 API 调用)", hasDesc === true); +} + +// ── W. 取消息出问题必须「响」——分类 + 抛异常,两头都不能省 ── +{ + const e1 = reportFetchProblem({ error: true, watermarkFound: true, watermark: "om_x" }); + check("W lark-cli 失败 → 判为 fetch-error", e1.ok === false && e1.reason === "fetch-error", JSON.stringify(e1)); + check("W 失败原因含排查提示(token/网络/权限)", /token|网络|权限/.test(e1.message || ""), e1.message); + + const e2 = reportFetchProblem({ error: false, watermarkFound: false, watermark: "om_x" }); + check("W 水位不在页内 → 判为 no-watermark", e2.ok === false && e2.reason === "no-watermark", JSON.stringify(e2)); + check("W 原因里带上水位 id 方便定位", (e2.message || "").includes("om_x"), e2.message); + + const e3 = reportFetchProblem({ error: false, watermarkFound: true, watermark: "om_x" }); + check("W 一切正常 → ok", e3.ok === true && e3.reason === "ok"); + + // 关键:pollOnce 不能靠"记得检查返回值"来保证响,必须靠抛异常。 + // 下面直接打真实的 collectFresh(注入假 fetch),不内联模拟逻辑。 + const mk = (msgs, err = false) => async () => ({ messages: msgs, caughtUp: true, error: err }); + // desc 页 = 最新在前。om_b 比 om_a 新,所以 om_a 是水位、水位后面有 om_b。 + const two = [{ message_id: "om_b" }, { message_id: "om_a" }]; + + let threw1 = null; + try { await collectFresh("c", "om_wm", { fetch: mk([]) }); } + catch (e) { threw1 = e; } + check("W 真 collectFresh:空页 → 抛异常(不是静默返回空)", !!threw1, `got ${threw1}`); + check("W 空页即使 error 标志没置,也判为取消息失败", + threw1?.reason === "fetch-error", String(threw1?.reason)); + check("W 异常 message 含排查提示", /token|网络|权限/.test(threw1?.message || ""), threw1?.message); + + let threw1b = null; + try { await collectFresh("c", "om_wm", { fetch: mk([], true) }); } + catch (e) { threw1b = e; } + check("W fetch 明确报错 → 同样是 fetch-error", threw1b?.reason === "fetch-error", String(threw1b?.reason)); + + let threw2 = null; + try { await collectFresh("c", "om_missing", { fetch: mk(two) }); } + catch (e) { threw2 = e; } + check("W 真 collectFresh:水位不在页内 → 抛异常", + !!threw2 && threw2?.reason === "no-watermark", String(threw2)); + check("W 异常 message 带水位 id", (threw2?.message || "").includes("om_missing"), threw2?.message); + + const okRes = await collectFresh("c", "om_a", { fetch: mk(two) }); + check("W 真 collectFresh:正常 → 返回新消息且不抛", + okRes.fresh.map(m => m.message_id).join(",") === "om_b" && okRes.caughtUp === true, + JSON.stringify(okRes.fresh)); +} + +console.log(results.join("\n")); +console.log(`\n${pass} passed, ${fail} failed, ${skipped} skipped`); +process.exit(fail ? 1 : 0); diff --git a/plugins/antianqi/mcode-feishu-bridge/skills/SKILL.md b/plugins/antianqi/mcode-feishu-bridge/skills/SKILL.md new file mode 100644 index 00000000..dc9e94d7 --- /dev/null +++ b/plugins/antianqi/mcode-feishu-bridge/skills/SKILL.md @@ -0,0 +1,176 @@ +--- +name: mcode-feishu-bridge +description: Drive MiniMax Code remotely from a Lark/Feishu conversation, so the user can send a task from their phone and get the result back in the same message. Use when the user wants to control mcode from Feishu or Lark, asks to set up, start, stop, inspect or troubleshoot a Feishu-driven mcode bridge, wants to know whether the bridge is running, or reports that a Feishu message got no reply. Requires the lark-cli and mcode CLIs; ships no credentials. +--- + +# mcode-feishu-bridge + +Runs a small polling daemon that turns a Lark/Feishu message into a local `mcode` +turn and writes the answer back into that same message. + +## What the user gets + +One Feishu message per task, updated in place: + +1. a placeholder appears within a poll interval, +2. it flips to `⚙️ calling \`write\`…` while tools run, +3. it ends as the final answer plus a `⏱ 12.3s · 1335 tokens · 2 tool calls` footer. + +Because every intermediate state is an edit of the same message, a task that makes +twelve tool calls still produces one bubble, not twelve. + +## Before starting anything + +> **Do not confuse this with the built-in `lark-tools` Skill.** `lark-tools` drives +> Feishu *from* mcode — documents, calendar, Base, mail, approvals. This Plugin is +> the opposite direction: a Feishu message becomes a local `mcode` turn. The +> desktop runtime also has a Feishu channel, which is the same reverse direction but +> needs the Mavis desktop app running and a Feishu app bound to it. Never present +> this Plugin as a replacement for `lark-tools`; they do different jobs and can run +> side by side. + +The bridge has **no credentials of its own**. It uses `lark-cli`, which must already +be configured for the user. Check first: + +```bash +lark-cli auth status +``` + +Two things must both be true, and they are the two that actually go wrong: + +- `identities.bot.status` is `ready` — the bridge **reads with the user identity but + writes with the bot identity**, because only a bot may edit a message it sent, and + a bot cannot see a p2p conversation's history. +- `identities.user.status` is `ready` and the token is `valid`. + +The app needs at least `im:message:readonly` and `im:resource` for the user side, and +`im:message` plus `im:message:update` for the bot side. `offline_access` matters most +in practice: without it the user token stops working after a couple of hours and the +bridge goes quiet with no error at the time it happens. + +If `lark-cli` is not signed in, do not attempt to sign in for the user. Tell them to +run `lark-cli config init` and `lark-cli auth login` themselves, and stop there. + +`mcode` must be installed and runnable. The bridge discovers both CLIs from `PATH` +and the install layout; it never hardcodes a path. + +## Configuration + +There is deliberately **no default chat id**. A chat id is a private identifier, so +the user has to supply one. The bridge accepts it three ways, in priority order: + +```bash +# 1. flag +node scripts/mcode-feishu-bridge.mjs --watch --chat oc_xxxxxxxxxxxxxxxx + +# 2. environment +MCODE_FEISHU_CHAT=oc_xxxxxxxxxxxxxxxx node scripts/mcode-feishu-bridge.mjs --watch + +# 3. config file in the data directory +# /config.json -> { "chatId": "oc_xxxxxxxxxxxxxxxx" } +``` + +The data directory is `$PLUGIN_DATA` when the runtime provides it, otherwise +`$MCODE_FEISHU_BRIDGE_DATA`, otherwise `~/.mcode-feishu-bridge`. + +## Operating the bridge + +```bash +# one pass, then exit — use this to check on a single message +node scripts/mcode-feishu-bridge.mjs --chat + +# long-running watcher +node scripts/mcode-feishu-bridge.mjs --watch --interval 2000 --chat + +# stop the running watcher (kills the mcode process tree too) +node scripts/mcode-feishu-bridge.mjs --stop +``` + +Only `--watch` takes the single-instance lock, so one-shot runs never block each +other. If a second watcher is refused with exit code 3, the first one is healthy; +use `--stop` rather than deleting the lock file. + +Useful flags: + +| flag | default | meaning | +|---|---|---| +| `--chat ` | none | conversation to watch | +| `--watch` | off | keep polling instead of one pass | +| `--interval ` | 3000 | poll interval | +| `--timeout ` | 600000 | hard ceiling on one mcode turn | +| `--log ` | `/bridge.log` in watch mode | also append output to a file | +| `--stop` | — | stop the running watcher | + +## Behaviour worth knowing before you debug it + +- **The bridge answers only what mcode can do.** It never fabricates a result. If + mcode fails, the failure text is what the user sees. +- **mcode runs with `--permission full`.** This is deliberate: a phone is a bad + place to answer approval prompts. Tell the user this plainly, because a + Feishu message can therefore cause arbitrary local changes. Suggest pinning the + bridge to a dedicated workspace directory if that trade is not acceptable. +- **One turn at a time per chat.** A second message is processed after the first + finishes, not concurrently. +- **A hung mcode is killed, not waited on.** After `--timeout` the process tree is + terminated and the chat gets an explicit timeout notice naming how many tool + calls had already run. + +## If the user says a message got no reply + +Work through this order; the first four are the ones that actually happen. + +1. **Read the log.** The bridge logs to `/bridge.log` and the file is + readable while it runs. An empty log plus a silent chat means the watcher is not + running at all — start it. + - No `✓` line for that message means it was never picked up. + - `✗ fetch messages failed` means the **user** identity failed. Usually an expired + token, or `offline_access` was never granted. Have the user re-run + `lark-cli auth login`. + - The answer is a failure notice even though the message was picked up: the **bot** + identity cannot send, so even the fallback reply failed. + - `✗ the watermark is not on this page` means the bridge fell too far behind; it + escalates to full pagination on its own, so this line means that also failed. +2. **Check both identities.** `lark-cli auth status`. `identities.user` covers + reading and attachments; `identities.bot` covers sending, replying and editing. + They fail independently and the symptoms differ. +3. **Check the process and the lock.** `bridge.lock` in the data directory holds a + pid. A lock whose pid no longer exists is stale and is reclaimed automatically. +4. **Check the watermark.** `state.json` records the last processed message per + chat. If it is ahead of the user's message, the message was already handled. +5. **Run the tests.** `node scripts/mcode-feishu-bridge.test.mjs` is self-contained + and prints a pass/fail summary. It is the fastest way to tell whether the host + is broken or the conversation simply outgrew one page. + +`README.md` has a symptom-to-cause table covering the permission and identity +failures, which are the ones users cannot diagnose from the log alone. + +## Tests + +```bash +node scripts/mcode-feishu-bridge.test.mjs +``` + +84 assertions covering edit ordering, delivery backoff and give-up, the mcode hard +timeout, the single-instance lock, message paging, and a static check that the +source never routes a child process through a shell. The suite imports the real +module rather than a copy, so a green run says something about the shipped code. + +Two cases (`O` and `V`) read this plugin's own source. If you add a spawn call, +keep `shell: false`; case `O` fails otherwise. + +## Running it unattended + +The Plugin ships no autostart mechanism, deliberately. If the user wants it always +on, put it under a supervisor they already run, and warn them about one specific +trap: **do not start a long-lived child with std handles inherited from a parent that +waits on the pipe.** The caller blocks until the child closes it, which is forever. +Start it detached, or let the bridge write its own log with `--log`. + +## Network destinations + +- `open.feishu.cn` / `open.larksuite.com`, through `lark-cli`, for message read, + message send, message edit, and attachment download. Required. +- The Feishu Open Platform host, through `lark-cli`, for token refresh. Required. + +No other destination is contacted. See `README.md` for the full data-flow +disclosure. diff --git a/plugins/antianqi/mcode-feishu-bridge/skills/mcode-feishu-bridge/SKILL.md b/plugins/antianqi/mcode-feishu-bridge/skills/mcode-feishu-bridge/SKILL.md new file mode 100644 index 00000000..dc9e94d7 --- /dev/null +++ b/plugins/antianqi/mcode-feishu-bridge/skills/mcode-feishu-bridge/SKILL.md @@ -0,0 +1,176 @@ +--- +name: mcode-feishu-bridge +description: Drive MiniMax Code remotely from a Lark/Feishu conversation, so the user can send a task from their phone and get the result back in the same message. Use when the user wants to control mcode from Feishu or Lark, asks to set up, start, stop, inspect or troubleshoot a Feishu-driven mcode bridge, wants to know whether the bridge is running, or reports that a Feishu message got no reply. Requires the lark-cli and mcode CLIs; ships no credentials. +--- + +# mcode-feishu-bridge + +Runs a small polling daemon that turns a Lark/Feishu message into a local `mcode` +turn and writes the answer back into that same message. + +## What the user gets + +One Feishu message per task, updated in place: + +1. a placeholder appears within a poll interval, +2. it flips to `⚙️ calling \`write\`…` while tools run, +3. it ends as the final answer plus a `⏱ 12.3s · 1335 tokens · 2 tool calls` footer. + +Because every intermediate state is an edit of the same message, a task that makes +twelve tool calls still produces one bubble, not twelve. + +## Before starting anything + +> **Do not confuse this with the built-in `lark-tools` Skill.** `lark-tools` drives +> Feishu *from* mcode — documents, calendar, Base, mail, approvals. This Plugin is +> the opposite direction: a Feishu message becomes a local `mcode` turn. The +> desktop runtime also has a Feishu channel, which is the same reverse direction but +> needs the Mavis desktop app running and a Feishu app bound to it. Never present +> this Plugin as a replacement for `lark-tools`; they do different jobs and can run +> side by side. + +The bridge has **no credentials of its own**. It uses `lark-cli`, which must already +be configured for the user. Check first: + +```bash +lark-cli auth status +``` + +Two things must both be true, and they are the two that actually go wrong: + +- `identities.bot.status` is `ready` — the bridge **reads with the user identity but + writes with the bot identity**, because only a bot may edit a message it sent, and + a bot cannot see a p2p conversation's history. +- `identities.user.status` is `ready` and the token is `valid`. + +The app needs at least `im:message:readonly` and `im:resource` for the user side, and +`im:message` plus `im:message:update` for the bot side. `offline_access` matters most +in practice: without it the user token stops working after a couple of hours and the +bridge goes quiet with no error at the time it happens. + +If `lark-cli` is not signed in, do not attempt to sign in for the user. Tell them to +run `lark-cli config init` and `lark-cli auth login` themselves, and stop there. + +`mcode` must be installed and runnable. The bridge discovers both CLIs from `PATH` +and the install layout; it never hardcodes a path. + +## Configuration + +There is deliberately **no default chat id**. A chat id is a private identifier, so +the user has to supply one. The bridge accepts it three ways, in priority order: + +```bash +# 1. flag +node scripts/mcode-feishu-bridge.mjs --watch --chat oc_xxxxxxxxxxxxxxxx + +# 2. environment +MCODE_FEISHU_CHAT=oc_xxxxxxxxxxxxxxxx node scripts/mcode-feishu-bridge.mjs --watch + +# 3. config file in the data directory +# /config.json -> { "chatId": "oc_xxxxxxxxxxxxxxxx" } +``` + +The data directory is `$PLUGIN_DATA` when the runtime provides it, otherwise +`$MCODE_FEISHU_BRIDGE_DATA`, otherwise `~/.mcode-feishu-bridge`. + +## Operating the bridge + +```bash +# one pass, then exit — use this to check on a single message +node scripts/mcode-feishu-bridge.mjs --chat + +# long-running watcher +node scripts/mcode-feishu-bridge.mjs --watch --interval 2000 --chat + +# stop the running watcher (kills the mcode process tree too) +node scripts/mcode-feishu-bridge.mjs --stop +``` + +Only `--watch` takes the single-instance lock, so one-shot runs never block each +other. If a second watcher is refused with exit code 3, the first one is healthy; +use `--stop` rather than deleting the lock file. + +Useful flags: + +| flag | default | meaning | +|---|---|---| +| `--chat ` | none | conversation to watch | +| `--watch` | off | keep polling instead of one pass | +| `--interval ` | 3000 | poll interval | +| `--timeout ` | 600000 | hard ceiling on one mcode turn | +| `--log ` | `/bridge.log` in watch mode | also append output to a file | +| `--stop` | — | stop the running watcher | + +## Behaviour worth knowing before you debug it + +- **The bridge answers only what mcode can do.** It never fabricates a result. If + mcode fails, the failure text is what the user sees. +- **mcode runs with `--permission full`.** This is deliberate: a phone is a bad + place to answer approval prompts. Tell the user this plainly, because a + Feishu message can therefore cause arbitrary local changes. Suggest pinning the + bridge to a dedicated workspace directory if that trade is not acceptable. +- **One turn at a time per chat.** A second message is processed after the first + finishes, not concurrently. +- **A hung mcode is killed, not waited on.** After `--timeout` the process tree is + terminated and the chat gets an explicit timeout notice naming how many tool + calls had already run. + +## If the user says a message got no reply + +Work through this order; the first four are the ones that actually happen. + +1. **Read the log.** The bridge logs to `/bridge.log` and the file is + readable while it runs. An empty log plus a silent chat means the watcher is not + running at all — start it. + - No `✓` line for that message means it was never picked up. + - `✗ fetch messages failed` means the **user** identity failed. Usually an expired + token, or `offline_access` was never granted. Have the user re-run + `lark-cli auth login`. + - The answer is a failure notice even though the message was picked up: the **bot** + identity cannot send, so even the fallback reply failed. + - `✗ the watermark is not on this page` means the bridge fell too far behind; it + escalates to full pagination on its own, so this line means that also failed. +2. **Check both identities.** `lark-cli auth status`. `identities.user` covers + reading and attachments; `identities.bot` covers sending, replying and editing. + They fail independently and the symptoms differ. +3. **Check the process and the lock.** `bridge.lock` in the data directory holds a + pid. A lock whose pid no longer exists is stale and is reclaimed automatically. +4. **Check the watermark.** `state.json` records the last processed message per + chat. If it is ahead of the user's message, the message was already handled. +5. **Run the tests.** `node scripts/mcode-feishu-bridge.test.mjs` is self-contained + and prints a pass/fail summary. It is the fastest way to tell whether the host + is broken or the conversation simply outgrew one page. + +`README.md` has a symptom-to-cause table covering the permission and identity +failures, which are the ones users cannot diagnose from the log alone. + +## Tests + +```bash +node scripts/mcode-feishu-bridge.test.mjs +``` + +84 assertions covering edit ordering, delivery backoff and give-up, the mcode hard +timeout, the single-instance lock, message paging, and a static check that the +source never routes a child process through a shell. The suite imports the real +module rather than a copy, so a green run says something about the shipped code. + +Two cases (`O` and `V`) read this plugin's own source. If you add a spawn call, +keep `shell: false`; case `O` fails otherwise. + +## Running it unattended + +The Plugin ships no autostart mechanism, deliberately. If the user wants it always +on, put it under a supervisor they already run, and warn them about one specific +trap: **do not start a long-lived child with std handles inherited from a parent that +waits on the pipe.** The caller blocks until the child closes it, which is forever. +Start it detached, or let the bridge write its own log with `--log`. + +## Network destinations + +- `open.feishu.cn` / `open.larksuite.com`, through `lark-cli`, for message read, + message send, message edit, and attachment download. Required. +- The Feishu Open Platform host, through `lark-cli`, for token refresh. Required. + +No other destination is contacted. See `README.md` for the full data-flow +disclosure.