Skip to content

Convert byte column offsets to characters - #26

Open
dpc00 wants to merge 1 commit into
SublimeLinter:masterfrom
dpc00:utf-columns
Open

dpc00 wants to merge 1 commit into
SublimeLinter:masterfrom
dpc00:utf-columns

Conversation

@dpc00

@dpc00 dpc00 commented Sep 29, 2026

Copy link
Copy Markdown

Fixes #25.

cppcheck reports the column ({column}) as a byte offset into the UTF-8 line, so every non-ASCII character before the error on the same line moved the highlight to the right, past the end of the line for errors near its end. Sublime counts characters.

reposition_match now converts the byte offset to a character offset using the source line (vv.select_line(line)), like the fix in SublimeLinter-xmllint. Both linters (cppcheck and cppcheck++) get it from a small abstract base class. ASCII lines are unchanged.

Checked with cppcheck 2.21.0 in Sublime Text 4215 (this branch loaded as SublimeLinter-cppcheck), on the file from the issue: with an e with acute accent and an emoji before them, the array error is now on the [ (start 18, it was 22) and the unread-variable error on the = (start 23, it was 27, i.e. on the line break).

Tests: tests/test_columns.py uses the real output of cppcheck for that file (source and output are in the test), for both linter classes, plus the converter on its own (multi-byte characters, an offset inside a character, past the end); 4 tests pass in Sublime Text 4215, flake8 with the repo config is clean.

🤖 Generated with Claude Code

https://claude.ai/code/session_01WtRdpWmSGFACg8HR3Vpk3d

cppcheck reports the column as a byte offset into the UTF-8 line, so
every non-ASCII character before the error moved the highlight to the
right (and past the end of the line for errors near its end). Sublime
counts characters.

Convert the column in `reposition_match` using the source line, in a
small abstract base class shared by both linters.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WtRdpWmSGFACg8HR3Vpk3d
@kaste

kaste commented Sep 30, 2026

Copy link
Copy Markdown
Member

Same as tslint, regression test should go through process_match

E.g.

   class TestRegex(unittest.TestCase):
       def assertMatch(self, string, expected, source):
           linter = Linter(sublime.View(0), {})
           match = next(linter.find_errors(string))

           # Use the supplied source as the main buffer, not an external file.
           match['filename'] = None
           actual = linter.process_match(match, VirtualView(source))
           self.assertIsNotNone(actual)

           # Ignore information we don't want to write down in the examples.
           self.assertEqual({k: actual[k] for k in expected}, expected)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

Error column is wrong when non-ASCII characters come earlier on the line (cppcheck counts bytes)

2 participants