You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Fixes gaps in vec merge and merge_parquet around collection metadata vs. column data, and lets vec merge use DuckDB. Updates to follow vecorel/specification#8
vec merge merges local GeoParquet files that are all in the target CRS with DuckDB, without loading them into memory; it falls back to the in-memory merge otherwise. New --engine option, --crs first keeps the first dataset's CRS (default remains EPSG:4326).
Breaking:vec merge keeps all properties by default; -i restricts to core + given properties, -e removes any property. Warns if required properties are dropped.
Schemas of a collection that occurs in multiple datasets are united instead of overwritten.
Array/object constants are hydrated correctly: they were spread over rows, dropped, or failed.
A missing collection is filled from a single schemas entry.
Collection-only properties that differ between datasets are removed with a warning.
merge_parquet:
hydrates constants with proper types (dates, binary, arrays, objects) and NaN as null
checks and suffixes ids per collection, and checks required properties per collection
recomputes the bbox, which was empty for GeoParquet 1.0 parts
All-null columns are no longer dehydrated, which wrote NaN into the collection metadata or failed for nullable integers.
get_pyarrow_type no longer mutates schemas with patternProperties.
Validation checks that every feature has a collection listed in schemas, checks required properties per collection, and checks the schemas of all collections, not just the first.
CHANGELOG restructured into Keep a Changelog categories.
Geometry exclusion behaves inconsistently across engines
vecorel_cli/conversion/duckdb.py:658
properties can omit geometry after --exclude geometry, but this restores it only for the DuckDB path. The GeoPandas path receives the exclusion unchanged and will either fail to write a spatial file or omit the geometry, while DuckDB silently keeps it. Reject geometry exclusion or normalize the property set before engine selection so both engines behave consistently.
Unhashable collection values abort validation
vecorel_cli/validation/geoparquet.py:145
A malformed non-scalar collection value makes unique()/set(...) raise TypeError: unhashable type before the validator can report the column's invalid type. Validation should handle arbitrary input values and record them as unknown rather than aborting.
Removed collection-only properties are incorrectly treated as present
vecorel_cli/vecorel/ops.py:84
This removes every collection-only property from missing, even when filtering removed that property from the merged collection. Excluding a required collection-only property declared by a retained external schema therefore produces an invalid file without the promised warning. Only treat collection-only properties that are still present in collection as satisfied.
The reason will be displayed to describe this comment to others. Learn more.
Review of what this adds on top of #55. 1 and 2 are reproduced with small scripts; the rest is from reading the code.
encoding/base.py:130: --engine geopandas fails on an array constant without a schema (tags=["x","y"] in both parts: ArrowTypeError); with a schema it stays a column on every row. DuckDB handles both. Suggest hydrating only the keys the merged collection drops, as merge_parquet does.
conversion/duckdb.py:405: results now depend on the engine. A null in a required property fails on DuckDB (with converter advice, and "collection" IS NULL repeated per group) but merges on geopandas; empty geometries are dropped on DuckDB only.
conversion/duckdb.py:83: _constants_table catches every exception and falls back to an inferred type, so an overflow or unparsable date is written with the wrong type and no warning.
vecorel/schemas.py:339: two versions of one extension (e.g. fiboa v0.2 and v0.3) are kept side by side in one collection; differing Vecorel versions do raise.
conversion/duckdb.py:420: the id check and suffixing use (collection, id) even with one collection, which slows the large-conversion path; the composite key is only needed with several collections.
merge.py:140: with -e, a non-GeoParquet input is read in full once for its column names and again for the merge.
conversion/duckdb.py:399: the null condition still quotes "{target}" by hand instead of _sql_name.
merge.py:153: get_duckdb_blocker duplicates the CRS check of _common_crs; two copies can disagree.
vecorel/ops.py:81: with -i, a required collection-only property is dropped without a warning.
vecorel/ops.py:29: rows with a null collection and no collection in the metadata now make create-geoparquet and the in-memory merge raise; merge_parquet swallows the same error and fails later.
* Merge: hydrate only what the parts disagree on (review #1)
* Merge: keep rows as they are with DuckDB, like in memory (review #2)
* Report constants that don't fit their schema type (review #3)
* Merge: reject two versions of one schema in a collection (review #4)
* DuckDB: key ids by collection only with several collections (review #5)
* Merge: apply excludes to the loaded data in memory (review #6)
* DuckDB: quote required property names with _sql_name (review #7)
* Share the CRS comparison of merge and DuckDB (review #8)
* Merge: warn when includes drop a required collection-only property (review #9)
* Merge: accept a collection that can't be determined (review #10)
* Add strict mode to vec merge, minor improvements and bug fixes (#60)
Co-authored-by: Matthias Mohr <m.mohr@moregeo.it>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #55.
Fixes gaps in
vec mergeandmerge_parquetaround collection metadata vs. column data, and letsvec mergeuse DuckDB. Updates to follow vecorel/specification#8vec mergemerges local GeoParquet files that are all in the target CRS with DuckDB, without loading them into memory; it falls back to the in-memory merge otherwise. New--engineoption,--crs firstkeeps the first dataset's CRS (default remains EPSG:4326).vec mergekeeps all properties by default;-irestricts to core + given properties,-eremoves any property. Warns if required properties are dropped.collectionis filled from a singleschemasentry.merge_parquet:get_pyarrow_typeno longer mutates schemas withpatternProperties.schemas, checks required properties per collection, and checks the schemas of all collections, not just the first.