feat(analytics): feed Lichess PGN dumps into Spark batch jobs
Build & Test (NowChessSystems) TeamCity build failed
Build & Test (NowChessSystems) TeamCity build failed
Add GameSource: normalises game records into a shared schema and selects backend via NOWCHESS_PGN_PATH. Unset = PostgreSQL game_records (unchanged); set = a Lichess PGN dump (file or http(s) URL). - Parse Lichess PGN with Spark SQL string functions only (no UDFs). - URLs fetched once via SparkContext.addFile, distributed to executors. - .pgn.zst decompressed in-process via zstd-jni, plain .pgn redistributed. - All four batch jobs read through GameSource and skip JDBC write-back in PGN mode (Parquet/CSV output only). Enables driving the analytics demo straight from https://database.lichess.org standard dumps. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
@@ -22,6 +22,9 @@
|
||||
// NOWCHESS_JDBC_URL (default: jdbc:postgresql://localhost:5432/nowchess)
|
||||
// NOWCHESS_DB_USER (default: nowchess)
|
||||
// NOWCHESS_DB_PASS (default: nowchess)
|
||||
// NOWCHESS_PGN_PATH (optional) — file or http(s) URL of a Lichess PGN dump (.pgn or .pgn.zst).
|
||||
// When set, all batch jobs read games from the dump instead of PostgreSQL and
|
||||
// skip JDBC write-back (Parquet/CSV output only). Demo data source.
|
||||
|
||||
plugins {
|
||||
id("scala")
|
||||
@@ -71,6 +74,10 @@ dependencies {
|
||||
|
||||
// PostgreSQL JDBC driver bundled so it is available on executor classpath.
|
||||
implementation("org.postgresql:postgresql:42.7.4")
|
||||
|
||||
// zstd-jni: decompress Lichess .pgn.zst dumps in-process. Provided at runtime by Spark
|
||||
// (it uses zstd-jni internally for shuffle/event-log compression), so compile-only here.
|
||||
compileOnly("com.github.luben:zstd-jni:1.5.6-9")
|
||||
}
|
||||
|
||||
application {
|
||||
|
||||
Reference in New Issue
Block a user