docs(integration): add Apache Spark, Apache Flink, and Trino guides - #155
Merged
Merged
Conversation
|
@majinghe is attempting to deploy a commit to the overtrue's projects Team on Vercel. A member of the Team first needs to authorize it. |
Add Spark, Flink, and Trino guides under the Data Analytics category in all five locales (en, zh, de, fr, ja), and wire them into the big-data meta.json and category landing pages. - spark.md: apache/spark:3.5.6 with hadoop-aws 3.3.4 via --packages, fs.s3a.* properties (path-style, plain HTTP), write/read Parquet at s3a://my-bucket/spark-demo/events. - flink.md: flink:1.20 session cluster with flink-s3-fs-hadoop copied into plugins/s3fs, s3.* properties via FLINK_PROPERTIES on both JobManager and TaskManager, batch filesystem sink write and filesystem source read. - trino.md: trinodb/trino:435 with the hive connector's file metastore located at s3://my-bucket/trino-metastore (metadata and data in RustFS), native S3 filesystem, schema/table/insert/select. - All three include a RustFS S3 Tables section linking /administration/data/s3-tables with the REST catalog connection values and the per-tool validation scope. Verified end to end on Ubuntu 24.04 against rustfs/rustfs-x86-musl: v2.3.1: Spark wrote and read back a 1000-row Parquet dataset; Flink wrote a 5-row CSV via a batch INSERT and read it back through a filesystem source; Trino created schema/table, inserted 5 rows, and selected them with metadata and data objects in RustFS.
majinghe
force-pushed
the
docs/spark-flink-trino
branch
from
September 21, 2026 02:32
0692579 to
de71186
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds three Data Analytics integration guides — Apache Spark, Apache Flink, and Trino — in all five locales (en, zh, de, fr, ja), wired into
big-data/meta.jsonand the category landing pages. All three link their upstream GitHub repositories in the product introductions and include a RustFS S3 Tables section pointing at S3 Tables with the REST catalog connection values and the per-tool validation scope from the support matrix.apache/spark:3.5.6+hadoop-aws:3.3.4via--packages,fs.s3a.*properties (path-style, plain HTTP), write/read Parquet ats3a://my-bucket/spark-demo/events. Troubleshooting covers the Hadoop version mismatch behindNumberFormatException: For input string: "60s".flink:1.20session cluster withflink-s3-fs-hadoopcopied from/opt/flink/opt/intoplugins/s3fs/,s3.*properties viaFLINK_PROPERTIESon both JobManager and TaskManager, batch-mode filesystem sink write + filesystem source read. Troubleshooting covers the missings3.*credentials, cross-network hostname resolution, and the "Stream closed" recovery quirk.trinodb/trino:435with the hive connector's file metastore located ats3://my-bucket/trino-metastore(metadata AND data in RustFS), native S3 filesystem (fs.native-s3.enabled), create schema/table, insert, select. Troubleshooting covers the per-version property names, the file-metastore location constraint, and CSV format limits.Verification
Validated end to end on Ubuntu 24.04 against
rustfs/rustfs-x86-musl:v2.3.1:s3a://my-bucket/spark-demo/events; read backROWS_READ_BACK: 1000;_SUCCESS+ snappy.parquet part objects confirmed viarc ls.flink-out/part-...-task-0-file-0(41 B) whose content is the exact 5 rows; filesystem source SELECT returned all rows.CREATE SCHEMA/CREATE TABLE/5-rowINSERT/SELECTall succeeded with metadata JSON and Parquet data object undertrino-metastore/demo/events/in RustFS.Console screenshots are light theme at 2× DPR, ≤300 KB: Chinese captures in
zh, English captures inen/de/fr/ja.npm run docs:checkpasses;npm run buildpasses (2735 pages); locale audit reports no errors for the new pages.