Repository navigation
[Fix] Include max_length and render options in GLM-5.3-Flash VL tokenize hash - #2145
Merged
jayhenry merged 1 commit intoOct 10, 2026
Conversation
…ize hash Glm53VLTokenizeFunction persists cache-path length predictions in a directory keyed by tokenize_fn.hash(), but _hash_str only encoded processor budgets and pack weights. Changing max_length, system_message, add_generation_prompt, enable_thinking, or reasoning_effort left the hash unchanged, so a later run reused old-config lengths: the packer budgets by old lengths while runtime truncates by new ones, and truncation landing inside a visual span degrades into silently packed fake samples. Encode the five parameters in _hash_str, reading max_length from the constructor local since super().__init__ stores it only afterwards. Add a regression test that classifies every __init__ parameter as hash-affecting or not and asserts hash() changes for each one; each variant rebuilds its baseline so the shared processor cache cannot mask the assertion.
ShilohYu
force-pushed
the
fix/glm53-vl-tokenize-hash-configs
branch
2 times, most recently
from
October 10, 2026 05:08
4082f72 to
15cbeb1
Compare
jayhenry
merged commit Oct 10, 2026
753f1e1
into
InternLM:feat/glm53flash-f1-vl-data
0 of 4 checks passed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This fixes stale on-disk length caches for GLM-5.3-Flash VL training when max_length or the chat render options change. The cache path predicts per-sample num_tokens under the current config and persists them in a directory keyed by tokenize_fn.hash(), but _hash_str only encoded processor budgets and pack weights. Changing max_length, system_message, add_generation_prompt, enable_thinking, or reasoning_effort left the hash unchanged, so a later run reused old-config lengths: the packer budgets by old lengths while runtime truncates by new ones, and truncation landing inside a visual span degrades into silently packed fake samples.
Encode the five parameters in _hash_str, reading max_length from the constructor local since super().init stores it only afterwards. Add a regression test that classifies every init parameter as hash-affecting or not and asserts hash() changes for each one; each variant rebuilds its baseline so the shared processor cache cannot mask the assertion. Unclassified new parameters fail the test.
Stack placement: targets
feat/glm53flash-f1-vl-data(#2109), so the fix can propagate into #2111. Note: existing on-disk length caches invalidate once on the next run, which is intended.This PR also carries a small unrelated commit that repairs the lint gate on the base branch: 3de92e7 added a selector argument to tilelang_dsa_topk_indices but left tilelang_indexer_topk_from_ranges referencing it undeclared, failing ruff F821 and mypy for every stacked PR. The parameter is threaded through, with a dispatch test for both selector values.
Validation: