Repository navigation
Conversation
`RelationalExpression` evaluated `IN` / `NOT IN` membership with Python `==`,
which is strict RDF *term* equality, while `=` and `!=` use `Identifier.eq`,
which is value equality. SPARQL 1.1 17.4.1.9 defines `IN` as exactly
(lhs = expression1) || (lhs = expression2) || ...
("The test is done with `=` operator, which tests for the same value"), and
17.4.1.10 defines `NOT IN` as its negation, so the two must never disagree.
They did. RDF 1.1 treats a simple literal and an `xsd:string`-typed literal as
the same term, and XSD type promotion makes `1` and `1.0` equal in value space,
so these all evaluated incorrectly:
"x"^^xsd:string IN ("x") -> false (should be true)
1 IN (1.0) -> false (should be true)
2 IN (<http://example/iri>, "str", 2.0) -> false (spec example: true)
while the equivalent `=` comparisons were correct.
The impact is silent and easy to misdiagnose: `FILTER(?v NOT IN (...))` passes
for a value that *is* in the list, so it wrongly admits rows rather than
dropping them. It was found in a SHACL `sh:sparql` constraint where literals
reach the query as `xsd:string`-typed (coerced by a JSON-LD context) while the
shape lists plain literals, which reported valid data as violating.
Use `Identifier.eq` -- the same comparison `=` already uses. It falls back to
`__eq__` for IRIs and blank nodes, so their behaviour is unchanged.
`Literal.eq` reports incomparable operands as `NotImplemented` rather than
raising, so that is normalised to a `SPARQLError`, exactly as the `=` path
below already does, and a comparison `TypeError` is normalised the same way.
This matters three times over: `NotImplemented` is truthy, so `IN` would
otherwise match anything once a list entry errored; evaluating it as a bool is
a `TypeError` on Python 3.14; and an escaping `TypeError` would abort the scan,
losing a later genuine match. Recording both as errors is what the spec
requires -- "Errors in comparisons cause the IN expression to raise an error if
the RDF term being tested is not found elsewhere in the list" -- which the
existing error accumulator already implements.
Adds parametrised regression tests covering all six normative examples from
17.4.1.9/17.4.1.10, plus value equality across plain, `xsd:string`-typed,
numeric, language-tagged and IRI operands, error-then-match ordering, and
genuine non-matches. The tests iterate the result rather than using
`len(list(...))`, which would swallow an evaluation `TypeError` via
`__length_hint__`.
Verified against the W3C SPARQL 1.0 and 1.1 conformance suites: pass/fail
counts are unchanged.
jdsika
force-pushed
the
fix/sparql-in-value-equality
branch
from
October 2, 2026 17:44
7b860d5 to
51c9b4f
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary of changes
RelationalExpressionevaluatesIN/NOT INmembership with Python==, which isstrict RDF term equality, while
=and!=useIdentifier.eq, which is valueequality. SPARQL 1.1 §17.4.1.9 defines
INas exactlyand §17.4.1.10 defines
NOT INasits negation. So the two must never disagree.
They do. RDF 1.1 treats a simple literal and an
xsd:string-typed literal as thesame term, and XSD type
promotion makes
1and1.0equal in value space. Onmaintoday:This also means
maingets one of the spec's own worked examples wrong:2 IN (<http://example/iri>, "str", 2.0)is specified astrue, but evaluates tofalse.Minimal reproduction:
The fix
Use
Identifier.eq, the same comparison=already uses. It falls back to__eq__forIRIs and blank nodes, so their behaviour is unchanged. The call is guarded by
isinstance(expr, Identifier)becauseexprcan legitimately be aSPARQLErrorinstance when the left-hand expression errors (
parserutils.Expr.evalreturns ratherthan raises), and that case keeps the previous
==path.Switching to
.eqalso changes the failure modes, so both are normalised exactly as the=path three lines below already does:NotImplemented—Literal.eqreports incomparable operands this way rather thanraising. It is truthy, so without a guard
INwould match anything as soon as onelist entry errored (
FILTER(?o IN (1/0))matched every row), and evaluating it as abool is a hard
TypeErroron Python 3.14.TypeError—Literal.eqraises this for e.g. two literals of the same unknowndatatype with different lexical forms. Left uncaught it would escape the loop, losing a
later genuine match (
"x"^^ex:custom IN ("y"^^ex:custom, "x"^^ex:custom)) and escapingevaluation entirely as a non-
SPARQLError.Recording both as errors rather than as "no match" is what the spec requires: "Errors in
comparisons cause the
INexpression to raise an error if the RDF term being tested isnot found elsewhere in the list". The pre-existing error accumulator already implements
that, so
2 IN (1/0, 2)istruewhile2 IN (3, 1/0)raises.Why this matters
The failure is silent and errs in the admitting direction:
FILTER(?v NOT IN (...))passes for a value that is in the list, so a guard intended to exclude rows lets them
through. It was found via a SHACL
sh:sparqlconstraint where literals reach the query asxsd:string-typed (coerced by a JSON-LD@context) while the shape lists plain literals;the constraint reported valid data as violating.
Backwards compatibility
A bug fix with no intended backwards-incompatible change.
IN/NOT INnow matchexactly the values
=already considered equal. Code relying on the previousterm-equality behaviour was relying on a deviation from the spec that already disagreed
with
=.W3C SPARQL 1.0 and 1.1 conformance suites: pass / skip / xfail / xpass counts are
unchanged.
Tests
Two parametrised tests. Together they cover all six normative examples from
§17.4.1.9 and §17.4.1.10 (
2 IN (1,2,3),2 IN (),2 IN (<http://example/iri>, "str", 2.0),2 IN (1/0, 2),2 IN (2, 1/0),2 IN (3, 1/0)), plus:xsd:string-typed literals, both directions1vs1.0)__eq__fallbackxsd:datevsxsd:dateTime, boolean vs numeric —so the fix cannot degrade into "always true"
Seven parametrisations fail on
mainand all pass with this change.The helper iterates the result rather than using
len(list(...)):Result.__len__isconsulted by
list()viaPyObject_LengthHint, which clears aTypeErrorraisedduring evaluation, so a crashing filter would silently look like "no rows" instead of
failing the test.
Checklist
the same change.
so maintainers can fix minor issues and keep your PR up to date.
Checks run locally:
black --check✓ ·ruff check✓ ·mypy --show-error-codes✓ ·pytest test/test_sparql test/test_w3c_spec/test_sparql{10,11}_w3c.py test/test_w3c_spec/test_sparql_rdflib.py→ 1341 passed, 0 failed.