Skip to content

docs(integration): add Apache Spark, Apache Flink, and Trino guides - #155

Merged
majinghe merged 1 commit into
rustfs:mainfrom
majinghe:docs/spark-flink-trino
Sep 21, 2026
Merged

majinghe merged 1 commit into
rustfs:mainfrom
majinghe:docs/spark-flink-trino

Conversation

@majinghe

Copy link
Copy Markdown
Collaborator

Summary

Adds three Data Analytics integration guides — Apache Spark, Apache Flink, and Trino — in all five locales (en, zh, de, fr, ja), wired into big-data/meta.json and the category landing pages. All three link their upstream GitHub repositories in the product introductions and include a RustFS S3 Tables section pointing at S3 Tables with the REST catalog connection values and the per-tool validation scope from the support matrix.

  • spark.md: apache/spark:3.5.6 + hadoop-aws:3.3.4 via --packages, fs.s3a.* properties (path-style, plain HTTP), write/read Parquet at s3a://my-bucket/spark-demo/events. Troubleshooting covers the Hadoop version mismatch behind NumberFormatException: For input string: "60s".
  • flink.md: flink:1.20 session cluster with flink-s3-fs-hadoop copied from /opt/flink/opt/ into plugins/s3fs/, s3.* properties via FLINK_PROPERTIES on both JobManager and TaskManager, batch-mode filesystem sink write + filesystem source read. Troubleshooting covers the missing s3.* credentials, cross-network hostname resolution, and the "Stream closed" recovery quirk.
  • trino.md: trinodb/trino:435 with the hive connector's file metastore located at s3://my-bucket/trino-metastore (metadata AND data in RustFS), native S3 filesystem (fs.native-s3.enabled), create schema/table, insert, select. Troubleshooting covers the per-version property names, the file-metastore location constraint, and CSV format limits.

Verification

Validated end to end on Ubuntu 24.04 against rustfs/rustfs-x86-musl:v2.3.1:

  • Spark: 1000-row Parquet dataset written to s3a://my-bucket/spark-demo/events; read back ROWS_READ_BACK: 1000; _SUCCESS + snappy.parquet part objects confirmed via rc ls.
  • Flink: batch INSERT wrote flink-out/part-...-task-0-file-0 (41 B) whose content is the exact 5 rows; filesystem source SELECT returned all rows.
  • Trino: CREATE SCHEMA/CREATE TABLE/5-row INSERT/SELECT all succeeded with metadata JSON and Parquet data object under trino-metastore/demo/events/ in RustFS.

Console screenshots are light theme at 2× DPR, ≤300 KB: Chinese captures in zh, English captures in en/de/fr/ja.

npm run docs:check passes; npm run build passes (2735 pages); locale audit reports no errors for the new pages.

@vercel

vercel Bot commented Sep 21, 2026

Copy link
Copy Markdown

@majinghe is attempting to deploy a commit to the overtrue's projects Team on Vercel.

A member of the Team first needs to authorize it.

Add Spark, Flink, and Trino guides under the Data Analytics category
in all five locales (en, zh, de, fr, ja), and wire them into the
big-data meta.json and category landing pages.

- spark.md: apache/spark:3.5.6 with hadoop-aws 3.3.4 via --packages,
  fs.s3a.* properties (path-style, plain HTTP), write/read Parquet at
  s3a://my-bucket/spark-demo/events.
- flink.md: flink:1.20 session cluster with flink-s3-fs-hadoop copied
  into plugins/s3fs, s3.* properties via FLINK_PROPERTIES on both
  JobManager and TaskManager, batch filesystem sink write and
  filesystem source read.
- trino.md: trinodb/trino:435 with the hive connector's file
  metastore located at s3://my-bucket/trino-metastore (metadata and
  data in RustFS), native S3 filesystem, schema/table/insert/select.
- All three include a RustFS S3 Tables section linking
  /administration/data/s3-tables with the REST catalog connection
  values and the per-tool validation scope.

Verified end to end on Ubuntu 24.04 against rustfs/rustfs-x86-musl:
v2.3.1: Spark wrote and read back a 1000-row Parquet dataset; Flink
wrote a 5-row CSV via a batch INSERT and read it back through a
filesystem source; Trino created schema/table, inserted 5 rows, and
selected them with metadata and data objects in RustFS.
@majinghe
majinghe force-pushed the docs/spark-flink-trino branch from 0692579 to de71186 Compare September 21, 2026 02:32
@majinghe
majinghe merged commit ca0b7c7 into rustfs:main Sep 21, 2026
1 of 2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant