Bring your own data
Point pragmatiq at real bank records — produce the parquet contract, validate, then run the same pipeline.
pragmatiq is an independent implementation inspired by the PRAGMA paper (arXiv 2604.08649) and is not affiliated with or endorsed by Revolut.
pragmatiq trains on the same parquet contract whether the data is synthetic or yours. The only new work is producing those four files; everything downstream is identical.
Produce the four files
Write events.parquet, profiles.parquet, and (optionally) transfers.parquet +
labels/*.parquet with the exact dtypes in the data contract.
The key idea: event fields and profile attributes are map<string,string> — pass raw
string values; the tokenizer infers numeric / categorical / text per field. Do not pre-bin or
pre-encode.
import pyarrow as pa, pyarrow.parquet as pq
events = pa.table({
"user_id": [...],
"ts": pa.array([...], type=pa.timestamp("us")),
"source": [...], # e.g. "transaction", "app", "trading"
"fields": [...], # list of {str: str}
})
pq.write_table(events, "data/mybank/events.parquet")Validate, then run the pipeline
pragmatiq validate data/mybank # fails loudly on dtype / integrity issues
pragmatiq tokenize data/mybank --out data/tok --n-workers 8
pragmatiq pretrain data/tok --name mybank --model-size small --config auto
pragmatiq probe data/tok --run runs/mybank --label data/mybank/labels/<task>.parquetAt 1M–26M records, --config auto sizes the batch/schedule for you, and the truncation caps
keep heavy-tailed histories tractable (see Pretrain).
Things to know
- Time zones: calendar features default to UTC; set
calendar_tz(e.g.Europe/London) if your timestamps are UTC instants but day/night and payday structure is local. - Unseen values at inference map to
[UNK]with a logged warning — never aKeyError. Refit the tokenizer when your value distribution shifts materially. - Forecast labels need an
eval_tsper user so histories are truncated before embedding (forecast, not hindcast). AML membership labels useuser_id,observed_through, andlabel.pragmatiq probe --staleness-window 6hadditionally drops the last hours before each eval point — the check that a lagging event feed is safe to serve from. - Long histories: per-event and profile token caps apply at encode time; the 6500-most-recent-events cap is applied when a batch is collated, after the eval-point cut, so shards keep the full history. Re-tokenize shards produced by 1.0.x.
Data never leaves your environment
Everything runs locally, on your GPU or CPU. The synthetic generator and
synth calibrate exist so you can develop and benchmark against a realistic book without
moving raw records.
AML GNN ablation
Run the four-arm GraphSAGE ablation over the transfer graph and read the result.
Run pragmatiq on Amazon SageMaker
Train the pragmatiq banking foundation model with an Amazon SageMaker training job and serve user embeddings from a NVIDIA Triton real-time endpoint — end to end, with pip install pragmatiq, on synthetic or your own data.