Skip to content

Python writer: Qwen2.5-Coder-0.5B on Core ML + live demo - #35

Open
Alex-Wengg wants to merge 2 commits into
mainfrom
feat/code-writer
Open

Alex-Wengg wants to merge 2 commits into
mainfrom
feat/code-writer

Conversation

@Alex-Wengg

@Alex-Wengg Alex-Wengg commented Oct 10, 2026 •

Copy link
Copy Markdown
Member

A 0.5B model on the Mac writes Python from plain-English requests and the code runs. Qwen2.5-Coder-0.5B-Instruct on Core ML (FluidInference/qwen2.5-coder-0.5b-coreml): the prompt runs on the Neural Engine, the writing on the GPU (same split as ShortReplyManager).

  • CodeWriterManager — loads the multifunction package (prefill 448 tokens on the ANE, decode with a 1,024-slot KV cache on the GPU), greedy decode, streams text after every token.
  • CodeWriterModelStore — checksummed download pinned to 48abc2c5.
  • CodeWriterDemo — ten MBPP tasks (one per Python feature) written live into an editor, then python3 runs each task's asserts; demo.sh adds macmon + the live log.
  • CodeWriterCheck — HumanEval through the Swift host, token-stream parity against a reference run.
M5 Pro, macOS 27
HumanEval pass@1 89/164 Core ML vs 90/164 PyTorch fp32 (142/164 identical token streams)
MBPP sanitized test split 119/257
Prompt / writing ~52 ms / 36–51 tok/s
Demo list 10/10 — hand-picked: Mbpp/12 failed and was replaced

Notes: the ANE carries ~1% of a task by time (prompt only; the stateful decoder runs on the GPU). Running the model's code needs python3 on the PATH. Conversion code lives outside this repo.

🤖 Generated with Claude Code

CodeWriterManager writes Python from a plain-English request with
Qwen2.5-Coder-0.5B-Instruct on Core ML (HF FluidInference/qwen2.5-coder-0.5b-coreml,
pinned @48abc2c5). One 944 MB multifunction package: `prefill` (448-token prompt,
Neural Engine, 1,451/1,456 ops) and `decode` (stateful 1,024-slot KV cache, GPU),
same split and K/V hand-off as ShortReplyManager. Greedy decode streams the text
after every token; extractCode pulls the first fenced block.

- CodeWriterModelStore: checksummed download of the pinned snapshot.
- CodeWriterCheck: replays HumanEval through the Swift host, compares token streams
  with a reference run, runs the tests with python3 (`hf` = download via the store).
- CodeWriterDemo: ten MBPP tasks (one per Python feature) written live into an
  editor, then each task's asserts run with python3; type-your-own bar; demo.sh
  opens macmon + the live log.

Numbers (M5 Pro, macOS 27): HumanEval pass@1 89/164 Core ML vs 90/164 PyTorch
fp32, 142/164 token streams identical; prefill ~52 ms, decode 36-51 tok/s. The demo
list is hand-picked (Mbpp/12 failed and was replaced by Mbpp/397); all ten pass.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant