Skip to content

Add lang-* and braille-* features to choose which rules include-zip embeds - #816

Open
trypsynth wants to merge 1 commit into
daisy:mainfrom
trypsynth:rules-subset
Open

trypsynth wants to merge 1 commit into
daisy:mainfrom
trypsynth:rules-subset

Conversation

@trypsynth

@trypsynth trypsynth commented Sep 24, 2026 •

Copy link
Copy Markdown

Why

I use MathCAT in Paperback, a screen reader friendly document reader, to read MathML in EPUB and HTML as AsciiMath. It works really well, thank you for it.

I only ever ask for one language (en) and one braille code (ASCIIMath). With include-zip, my binary carries all fifteen languages and all nine braille codes anyway, because build.rs stages the whole Rules tree into one rules.zip and shim_filesystem.rs embeds it. That is 857 KB of the binary, and about 790 KB of it is rules nothing in my program can ask for. The same weight lands in the iOS framework and the Android libraries, where app size is something users notice.

What this does

Every Rules/Languages and Rules/Braille directory gets a feature, and all of them are on by default:

default = ["rules-all"]
rules-all = ["lang-all", "braille-all"]
lang-all = ["lang-de", "lang-el", "lang-en", ...]
braille-all = ["braille-asciimath", "braille-cmu", ...]
lang-en = []
braille-asciimath = []
...

A consumer that wants a subset writes:

mathcat = { version = "...", default-features = false, features = ["include-zip", "lang-en", "braille-asciimath"] }

A default build is unchanged. I checked that the embedded rules.zip is byte for byte identical to the one from main. The features only affect what include-zip embeds. Without include-zip, the Rules directory is read from disk as before.

A build that subsets prints a cargo::warning naming what it kept, so nobody is surprised later by a missing language.

Keeping the lists up to date

build.rs compares the Rules/Languages and Rules/Braille directories with lang-all and braille-all in Cargo.toml. If a directory has no feature listed there, the build fails with:

Rules/Languages/xx has no feature: add `lang-xx = []` to Cargo.toml and list it in `lang-all`

This runs on every build, not only include-zip ones, so whoever adds a language finds out right away instead of it quietly dropping out of every build. The parse handles the one-entry-per-line form that cargo publish writes. I built the packaged .crate as a dependency to confirm that.

Fallbacks

The library still falls back by name, so the patch works around those fallbacks rather than changing them:

  • en is always embedded, because find_file and set_style_file fall back to it. It is about 78 KB compressed.
  • A subset with no braille code keeps UEB. Without any braille directory, set_rules_dir fails with Wasn't able to find/read MathCAT default language directory: Rules\Braille\, even for a speech-only caller.
  • If the BrailleCode default in prefs.yaml (Nemeth) was pruned, the staged copy points at a code that was kept. Without this, en plus ASCIIMath fails the same way.

One known rough edge remains. With only en embedded, set_preference("Language", "es") returns Ok(()) and then speaks English. Turning that into an error means changing the fallbacks themselves, which is better as its own change.

Copy as ASCIIMath / LaTeX

Applications that copy math as ASCIIMath or LaTeX, like the NVDA add-on, do it through the braille API. They need braille-asciimath / braille-latex. The comment above the features in Cargo.toml says so.

Tested

A small consumer, built with include-zip:

features result size (debug)
rules-all es speech "a partido por b", Nemeth ⠹⠁⠌⠃⠼ 9.41 MB
lang-en, braille-asciimath speech "eigh over b", ASCIIMath a/b 8.62 MB
lang-es only speech "a partido por b", UEB braille 8.66 MB

cargo clippy --all-targets --features include-zip is clean for the changed files. cargo test passes except shim_filesystem::tests::in_memory_filesystem_shim, which fails on Windows on main too.

@NSoiffer

Copy link
Copy Markdown
Collaborator

I agree that a feature flag is a better way to deal with this (probably 2 features, one for braille and one for language).

You are correct that MathCAT currently requires "en". I thought that was also true for "UEB". I suspect that both could be removed as defaults with some effort at the cost of generating errors instead of some sort of speech or braille.

You have to be careful about pruning the braille ASCIIMath and LaTeX because they are used for copy. If you don't allow copying in those formats, then it is probably a non issue, but testing is needed.

If no braille is needed, it is likely that all or most of braille.rs can be cfg'd out as can a little bit of speech.rs, thus shrinking the size of the binary some more.

@trypsynth

trypsynth commented Sep 25, 2026 •

Copy link
Copy Markdown
Author

Thanks for looking at it, and for the warning about copy, which I had not considered at all.

Features rather than environment variables: agreed, and two families is what I had in mind too, one per language and one per braille code. I will rewrite it that way.

One design question before I do, because cargo features are additive and "only English" is a subtractive statement. The way to express it is a default feature that carries everything:

[features]
default = ["rules-all"]
rules-all = ["lang-all", "braille-all"]
lang-en = []
braille-asciimath = []

Then every existing user is unaffected, because they get default, and a consumer who wants a subset writes default-features = false, features = ["include-zip", "lang-en", "braille-asciimath"]. Does that suit you, or would you rather the selection features were purely additive on top of a smaller default? The first keeps compatibility, the second is cleaner but is a breaking change.

On the fallbacks: I said something different in an earlier version of this comment and I have changed my mind, so please ignore that. Removing them is the better answer, and my own pull request description already contains the reason. With only en embedded, set_preference("Language", "es") returns Ok(()) and then speaks English, confidently and in the wrong language. That is worse than an error, and it is the kind of thing that costs somebody a long afternoon to work out.

So if you are willing to have find_file stop falling back by name, I would rather build on that than work around it. The distinction I would keep is between a full build and a subset build: a default build should behave exactly as it does today, and only a build that deliberately asked for a subset should get the error, since it is the one that knows what it left out. That way nobody's existing behaviour changes and the new failure mode is loud in exactly the case where silence is dangerous.

If you would rather not touch the fallbacks at all, the patch works without it: it force-keeps Languages/en whatever the caller asked for, and rewrites the staged prefs.yaml so BrailleCode defaults to a code the subset kept, which gets set_rules_dir returning Ok(()) with no library change. I am happy to ship that version. It is just not the one I would pick.

On ASCIIMath and LaTeX being used for copy: this is the part I am glad you raised. I had only thought about reading, so I would have shipped a feature set that silently broke copying for someone. Two things follow:

  1. Could you point me at the copy path? If it calls get_braille with a code the caller never sets as a preference, then a consumer pruning that code gets a runtime error from a feature list that looked reasonable, which is the worst shape for this. If that is how it works I would rather braille-asciimath and braille-latex be implied by something, or at minimum be loudly documented.
  2. Paperback itself is safe either way: we only ever read AsciiMath as text and never offer copy in another notation. So I cannot test the copy paths meaningfully, which is an argument for you or someone who uses them reviewing that part rather than me asserting it works.

On cfg'ing out braille.rs: that is clearly the bigger win, and I would be glad to do it as a follow-up once this lands. It is no use to Paperback, since AsciiMath reaches us through the braille API, but a speech-only consumer would get a lot more than the 771 KB this PR measures.

I will push the features version in the next few days and leave the environment-variable commit in the history for reference rather than in the branch.

@NSoiffer

Copy link
Copy Markdown
Collaborator

It is too bad that cargo features are either on or off and can't take values. So you are right that forces something like

[features]
default = ["rules-all"]

# Umbrella features
rules-all = ["lang-all", "braille-all"]
lang-all = ["lang-en", "lang-es", "lang-de", "lang-vi", "lang-zh", ...]
braille-all = ["braille-asciimath", "braille-ueb", "braille-vietnam", ...]

# Individual language flags
lang-en = []
lang-es = []
lang-de = []
lang-vi = []
lang-zh = []
....

# Individual braille flags
braille-asciimath = []
braille-ueb = []
braille-vietnam = []
...

It also means keeping thoses lists up-to-date which is error-prone. But I don't see a way around that.

As for copying: that is not in MathCAT directly. That is left up to the application. See https://github.com/daisy/MathCATForPython/blob/main/addon/globalPlugins/MathCAT/MathCAT.py#L394C1-L448C1 along with maybe some of the lines after it.

@trypsynth trypsynth changed the title Let a build pick which rules get embedded (MATHCAT_INCLUDE_LANGUAGES / MATHCAT_INCLUDE_BRAILLE) Add lang-* and braille-* features to choose which rules include-zip embeds Oct 1, 2026
@trypsynth

Copy link
Copy Markdown
Author

Thanks, that copy pointer was exactly what I needed. I've pushed the features version, with the layout you sketched, and updated the description.

On the lists being error-prone: build.rs now fails if any Rules/Languages or Rules/Braille directory isn't listed in lang-all / braille-all, and the error message names the line to add. So adding a language without its feature can't get past the first build.

Two smaller things came out of testing:

  • A subset with no braille code at all couldn't call set_rules_dir, so it keeps UEB the same way en is kept.
  • The default build's embedded rules.zip is byte for byte identical to main's, so existing users see no change.

I left the fallbacks themselves alone for now. Cfg'ing out braille.rs for speech-only builds can follow separately if you'd like it.

…mbeds

Every language and braille code has a feature, all on by default through rules-all, so existing builds embed exactly the same rules.zip. A build that turns off default features and names a subset embeds only those directories. en is always kept because the library falls back to it, a subset with no braille code keeps UEB so set_rules_dir still succeeds, and the staged prefs.yaml BrailleCode default is pointed at a kept code when Nemeth was pruned. build.rs fails if a Rules directory has no feature in lang-all or braille-all, so the lists cannot silently fall behind.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Triage

Development

Successfully merging this pull request may close these issues.

2 participants