Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@
- beartype is now `^0.22`. At 0.20.0 and below, the beartype import hook leaves `cli` a plain function instead of a group, so the command line either fails to import or runs `convert` whatever the arguments are. 0.20.1 is the first version that works.
- `in2lambda source add FILE` freezes a document: `in2lambda source add` converts .docx and .tex to markdown beside the file, and writes `FILE.draft.json` beside the file, holding the markdown's hash and every block in the markdown with the lines it spans, so that another tool can quote the source by line range. The draft is named after the source, so a folder holding a term's sheets holds one draft per sheet. `in2lambda source show` prints that markdown numbered with the block ids. Freezing a file that has changed since is refused unless `--start-over` discards the draft. Showing a file that has changed is refused as well, because its block ids would name lines they were not taken from. `in2lambda source add` needs pandoc and the `convert` extra, as `in2lambda convert` does. `in2lambda source show` reads the markdown the draft was frozen from, and needs neither.
- `in2lambda source add` writes the markdown of a converted .docx or .tex unwrapped, so that a paragraph is one line however long the paragraph is, and an inline `$ ... $` is never broken across two lines. `in2lambda source add` also moves every `$$ ... $$` that pandoc wrote on one line onto lines of its own, in a paragraph and in a list item, indented to the item's width where the maths is written in a list item. A `$$` that opens or closes in a table cell, in a block quote or in a code block is left unchanged, as is the maths that an unpaired `$$` elsewhere in the document, such as one in inline code, pairs with. `in2lambda validate` reports the maths left unchanged. pandoc's writer produces both forms, and Lambda Feedback renders neither, so `in2lambda validate` reported them against every converted sheet. The markdown beside a document frozen before this release differs, and so do its line ranges. `in2lambda source add --start-over` freezes the document again.
- `in2lambda source add` now drops the underline, highlight and small capitals a converted .docx holds, keeping the text each marked, so a field reads `**Question 2:**` in place of `**[Question 2:]{.underline}**`. Pandoc's docx reader writes those three formats as the bracketed spans `[text]{.underline}`, `[text]{.mark}` and `[text]{.smallcaps}`, which the PDF generator's pandoc turns into LaTeX commands its template does not define, so a set built from such a field failed to compile. Lambda Feedback's markdown renders none of the three. A bracketed span in a code block or in inline code is left unchanged. The markdown beside a .docx frozen before this release differs, and so do its line ranges where a span stood; `in2lambda source add --start-over` freezes the document again.
- A draft now holds a `log` of every command that changed it, and a `fields` map of the fields those commands wrote. Each field records the layer that wrote it (1 a spec, 2 a predicate, 3 a line range, 4 a literal), the source ranges it was copied from, whether it has been edited, and by whom. `in2lambda draft mark ignore BLOCK` is the first such command. `in2lambda draft replay` rebuilds the draft from the frozen markdown and the log, and refuses unless the draft it builds matches the draft on disk byte for byte. A draft written before this release holds no `log` and is refused as a draft in2lambda did not write; `in2lambda source add --start-over` freezes the document again.
- `in2lambda draft question add`, `in2lambda draft part add QUESTION` and `in2lambda draft question solution QUESTION` fill a draft in. Each takes `--text` to copy the wording out of the frozen source, as a block id such as `b3` or as lines such as `s10:14`. Each takes `--literal TEXT` for wording the source does not hold in a form the field can take, which records the field as edited and written by layer 4 in place of layer 3. in2lambda works the question and part numbers out from the fields already written, so a replay arrives at the same ids. `in2lambda draft split block BLOCK AT` cuts a block the parser made one of two things into `b3a` and `b3b`, so that each half can be quoted on its own. A command writing a field that is already written is refused, naming the field. A command quoting lines another field was taken from is refused, naming both fields.
- `in2lambda draft part solution PART --text RANGE` gives one part of a question its worked solution, for a sheet that writes a solution under each part rather than one solution answering the whole question. PART is written `q1.p2`. The command writes `q1.p2.solution` from the lines named, or from `--literal TEXT`, as `in2lambda draft question solution` writes a question's solution. A PART naming a question, or naming a part the draft has not written, is refused, naming the part. A sheet laid out this way can now be answered part by part by commands, which only a spec's `PartPartSolSol` and `PartSolPartSol` layouts could do before, so the warning `in2lambda validate` prints about a part nothing answers can be acted on.
Expand Down
96 changes: 91 additions & 5 deletions in2lambda/source/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -97,6 +97,12 @@ def _field_fault(field: Any) -> str:
The writer and the reader must agree. Pandoc's ``markdown`` writer emits fenced divs and
bracketed spans that a commonmark reader reads as ordinary text. ``commonmark_x`` also
covers the ``$...$`` maths and the ``{width=...}`` attributes a converted document holds.

``commonmark_x`` does write the bracketed spans a .docx holds, which
:func:`_spans_unwrapped` unwraps afterwards. The ``-bracketed_spans`` the writer takes
does not help: with ``raw_html`` on, which ``commonmark_x`` keeps, the writer falls back
to ``<u>Question 2:</u>`` and ``<span class="mark">oil</span>``, and turning ``raw_html``
off as well changes how a figure is written.
"""

_POSITION = re.compile(r"(?:[^@;]*@)?(\d+):\d+-(\d+):(\d+)")
Expand Down Expand Up @@ -422,6 +428,81 @@ def blocked(position: int) -> bool:
return "".join(written)


_ATTRIBUTE = r"""[.#][^\s{}]+|[\w-]+=(?:"[^"\n]*"|[^\s{}]+)"""
"""One attribute of a pandoc attribute list: a class, an id, or a key and its value."""

_SPAN = re.compile(rf"\[([^\[\]\n]*)\]\{{(?:{_ATTRIBUTE})(?: +(?:{_ATTRIBUTE}))*\}}")
"""A bracketed span as ``commonmark_x`` writes one, opened and closed on the one line.

The ``]{`` is what tells one from a link's ``](`` and from an image's ``){width=...}``.
The braces must hold an attribute list, because LaTeX writes brackets before braces as
well: ``$\\sqrt[3]{x + 1}$`` is a cube root, and dropping its braces would leave
``$\\sqrt3$``, which KaTeX renders and no check reports.
"""


def _spans_unwrapped(markdown: str) -> str:
r"""Markdown pandoc wrote, with the attributes of its bracketed spans dropped.

Pandoc's docx reader turns Word's underline into ``[Question 2:]{.underline}``, its
highlight into ``[oil]{.mark}`` and its small capitals into ``[Note:]{.smallcaps}``.
A field quoted out of markdown holding one of those is read back by the PDF
generator's pandoc as underline, highlight or small capitals, and written to LaTeX as
a command the generator's template does not define, so the set fails to compile.
Lambda Feedback's markdown renders none of the three, so the attribute is dropped and
the text it marked is kept. The reader emits a ``[text]{custom-style=...}`` span only
with its ``+styles`` extension, which the freeze does not enable; a document holding
one is unwrapped the same way.

Every rewrite stays within the line it began on, so no line range moves. A span
pandoc broke over two lines is left as written, and ``in2lambda validate`` reports the
field quoting it. A span in a code block or in inline code is left as written as well:
those are characters the document shows.

Examples:
>>> from in2lambda.source import _spans_unwrapped
>>> _spans_unwrapped("# [Hydraulic scale]{.underline}\n")
'# Hydraulic scale\n'
>>> _spans_unwrapped("**[Question 2:]{.underline}** joined by [oil]{.mark}.\n")
'**Question 2:** joined by oil.\n'
>>> _spans_unwrapped("[**[a]{.mark}**]{.underline}\n")
'**a**\n'
>>> _spans_unwrapped("[Note:]{.smallcaps} the oil is incompressible.\r\n")
'Note: the oil is incompressible.\r\n'
>>> _spans_unwrapped("Type `[a]{.mark}` first.\n")
'Type `[a]{.mark}` first.\n'
>>> _spans_unwrapped("Type this:\n\n [a]{.mark}\n")
'Type this:\n\n [a]{.mark}\n'
>>> _spans_unwrapped("::: {.solution}\nThe load is $F = pA$.\n:::\n")
'::: {.solution}\nThe load is $F = pA$.\n:::\n'
>>> _spans_unwrapped('![](figure.png){width="1in"}\n')
'![](figure.png){width="1in"}\n'
>>> _spans_unwrapped("The root is $\\sqrt[3]{x + 1}$.\n")
'The root is $\\sqrt[3]{x + 1}$.\n'
"""
if "\r\n" in markdown:
# Pandoc writes the line endings of whoever is running it, and the file on disk
# is hashed as it is written, so a Windows freeze stays a Windows file.
return _spans_unwrapped(markdown.replace("\r\n", "\n")).replace("\n", "\r\n")

verbatim = _verbatim_lines(markdown)

def unwrapped(match: re.Match[str]) -> str:
before = markdown[markdown.rfind("\n", 0, match.start()) + 1 : match.start()]
if markdown.count("\n", 0, match.start()) + 1 in verbatim or (
before.count("`") % 2
):
return match.group()
return match.group(1)

rewritten = _SPAN.sub(unwrapped, markdown)
if rewritten == markdown:
return markdown
# A span holding a span - `[**[a]{.mark}**]{.underline}` - unwraps from the inside,
# because the brackets of the outer one hold the brackets of the inner one.
return _spans_unwrapped(rewritten)


def _digest(data: bytes) -> str:
"""How a frozen markdown is named in its draft, so that a change to it is reported.

Expand Down Expand Up @@ -870,7 +951,9 @@ def add(
are written to ``FILE.draft.json``, so that code quoting a source by line range can
check that those lines still hold the text they held. A converted file is written
unwrapped: a paragraph is one line, however long, and each ``$$ ... $$`` is written
on lines of its own. Lambda Feedback renders display maths written that way.
on lines of its own. Lambda Feedback renders display maths written that way. The
underline, highlight and small capitals a .docx holds are dropped and the text they
marked is kept, because Lambda Feedback renders none of the three.

The files are numbered in the order they are given. A file already frozen into the
draft beside them is checked against the hash it was frozen at, and is not frozen
Expand Down Expand Up @@ -929,11 +1012,14 @@ def add(
raw, markdown = _source(path)
frozen_path = path
else:
# Unwrapped, and with the display maths blocked out, before anything is
# hashed: both are habits of pandoc's writer rather than anything the author
# did, and both are what a field quoting these lines would have to render.
# Unwrapped, with the display maths blocked out and the bracketed spans
# dropped, before anything is hashed: all three are habits of pandoc's writer
# rather than anything the author did, and all three are what a field quoting
# these lines would have to render.
markdown = _display_maths_blocked(
_pandoc(str(path), _MARKDOWN, "--wrap=none").decode("utf-8")
_spans_unwrapped(
_pandoc(str(path), _MARKDOWN, "--wrap=none").decode("utf-8")
)
)
raw = markdown.encode("utf-8")
frozen_path = path.with_suffix(".md")
Expand Down
9 changes: 9 additions & 0 deletions tests/fixtures/sources/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -62,3 +62,12 @@ image usually has no alt text, so this is also the ordinary case.

The image the docx embeds is referenced as `media/rId9.png`, which is not extracted - nothing
here reads the image, only the lines around it.

`bracketed_spans/source.docx` holds an underlined heading, a bold and underlined `Question 2:`
beside a highlighted word, and a run in small capitals. `make_source.py` beside it wrote that
document with python-docx, which is not a dependency of in2lambda: the document is committed and
the script records what it holds. Pandoc's docx reader turns the three formats into the bracketed
spans `[Hydraulic scale]{.underline}`, `[oil]{.mark}` and `[Note:]{.smallcaps}`, which the freeze
unwraps to the text alone. The test asserting that no fixture's frozen markdown holds a `]{`
looks for that rather than for `{.`, because `fenced_div` freezes to `::: {.solution}`, which is a
div the freeze keeps.
20 changes: 20 additions & 0 deletions tests/fixtures/sources/bracketed_spans/expected.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
[
{
"end": 1,
"id": "b1",
"start": 1,
"type": "heading"
},
{
"end": 3,
"id": "b2",
"start": 3,
"type": "paragraph"
},
{
"end": 5,
"id": "b3",
"start": 5,
"type": "paragraph"
}
]
36 changes: 36 additions & 0 deletions tests/fixtures/sources/bracketed_spans/make_source.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
"""Writes the ``source.docx`` beside this file, which was run by hand.

python-docx is not a dependency of in2lambda and is not added as one: the document is
committed, and this script records what it holds so that anyone can write it again. The
import stands inside the main block because the test suite imports every module under
``tests`` to collect its doctests.

Word's underline, highlight and small capitals are the three formats pandoc's docx
reader turns into a bracketed span, which `in2lambda source add` unwraps.
"""

if __name__ == "__main__":
from pathlib import Path

from docx import Document
from docx.enum.text import WD_COLOR_INDEX

document = Document()

heading = document.add_heading("", level=1)
heading.add_run("Hydraulic scale").underline = True

paragraph = document.add_paragraph()
label = paragraph.add_run("Question 2:")
label.bold = True
label.underline = True
paragraph.add_run(" A hydraulic scale has two pistons joined by ")
paragraph.add_run("oil").font.highlight_color = WD_COLOR_INDEX.YELLOW
paragraph.add_run(".")

note = document.add_paragraph()
note.add_run("Note: ").font.small_caps = True
note.add_run("the oil is incompressible.")

# Beside this file, so that the fixture is written wherever the script is run from.
document.save(Path(__file__).with_name("source.docx"))
Binary file not shown.
36 changes: 34 additions & 2 deletions tests/test_source.py
Original file line number Diff line number Diff line change
Expand Up @@ -14,15 +14,18 @@

import pytest
from click.testing import CliRunner
from conftest import SOURCES, SOURCES_DIR
from conftest import SOURCES, SOURCES_DIR, needs_compiler

from in2lambda.draft import _quoted
from in2lambda.main import cli
from in2lambda.validation import MathDelimiterError, math_delimiter_checker
from in2lambda.validation import MathDelimiterError, math_delimiter_checker, pdf

MARKDOWN = SOURCES_DIR / "markdown"
"""The case the tests below happen to use; what they check holds for any of them."""

SPAN_ATTRIBUTE = re.compile(r"\]\{[^}]*\}")
"""The attributes of a bracketed span, which no frozen markdown may hold."""


def _frozen(directory: Path) -> Path:
"""The document a folder holds, whatever format it is in."""
Expand Down Expand Up @@ -53,6 +56,35 @@ def test_source_add_finds_the_expected_blocks(folder: Path, tmp_path: Path) -> N
quoted = _quoted(draft, markdown, 1, block["start"], block["end"])
assert math_delimiter_checker(quoted) is MathDelimiterError.PASSED, quoted

# A bracketed span is underline, highlight or small capitals, which the PDF
# generator's pandoc writes as a LaTeX command its template does not define, so no
# freeze may leave one. A fenced div's `::: {.solution}` is not one of these.
assert not SPAN_ATTRIBUTE.search(markdown)


@needs_compiler
def test_a_frozen_docx_compiles_as_the_pdf_generator_does(tmp_path: Path) -> None:
"""A .docx holding underline, highlight and small capitals is what faulted at build.

The three sheets that reported it are private, so the fixture stands for them: every
block of it is quoted as a field would be and compiled as Lambda Feedback compiles a
set.
"""
shutil.copytree(SOURCES_DIR / "bracketed_spans", tmp_path, dirs_exist_ok=True)

result = CliRunner().invoke(cli, ["source", "add", str(tmp_path / "source.docx")])

assert result.exit_code == 0, result.output
draft = json.loads((tmp_path / "source.draft.json").read_text())
(source,) = draft["sources"]
markdown = (tmp_path / source["source"]).read_text()
fields = [
(block["id"], _quoted(draft, markdown, 1, block["start"], block["end"]))
for block in source["blocks"]
]

assert pdf.problems(fields, []) == []


def test_freezing_again_is_refused_once_the_source_has_changed(
tmp_path: Path, monkeypatch
Expand Down
Loading