Quickstart
Install pragmatiq and run the full pipeline — generate data, tokenize, pretrain, embed, and probe — on a GPU or a laptop CPU.
pragmatiq is an independent implementation inspired by the PRAGMA paper (arXiv 2604.08649) and is not affiliated with or endorsed by Revolut.
This runs the entire pragmatiq pipeline end to end on synthetic data. By the end you will have a trained model and a credit-risk probe score measured against a raw-count baseline — proof the embedding carries signal beyond trivial counts.
Requirements
Python 3.11+, torch 2.6+. A GPU is picked up automatically when visible (see Running on GPU); on a CPU the fast path below takes a few minutes in fp32. The commands use a virtual environment so the install is isolated.
Install
git clone https://github.com/dynamiq-ai/pragmatiq.git
cd pragmatiq
python3.11 -m venv .venv && source .venv/bin/activate
pip install -e ".[train]"The train extra brings in Lightning, which the pretraining stage runs on; the plain
install is the slim inference core (validate, tokenize, embed, probe with a trained
run). Other extras add focused tooling — .[serve] for ONNX/Triton export, .[aml]
for the transfer-graph GNN, .[text] for the frozen text embedder, .[tracking] for
experiment tracking, .[demo] for the Streamlit app, .[full] for all of them — see
Install for the full table.
Run the whole pipeline
pragmatiq quickstart --n-users 2000 --max-steps 80 # the fast path: a few minutes
pragmatiq quickstart # the reference run (50k users, 400 steps)from pragmatiq import api
result = api.quickstart(n_users=2000, max_steps=80)
print(result["message"])quickstart runs five stages in order:
- generate synthetic users and event histories,
- fit the key–value–time tokenizer,
- pretrain a nano masked-language model,
- embed users,
- probe a gradient-boosting credit-risk classifier against a raw-count baseline.
Read the result
The run prints a one-line summary like:
credit probe beats the raw-count baseline · run: runs/quickstartThe probe head is gradient boosting by default (HistGradientBoostingClassifier),
and the raw-count baseline uses the same classifier — so the gap reflects the
representation, not the model family. Both ROC-AUC and PR-AUC are reported (PR-AUC is
the honest headline on low-prevalence risk tasks).
Why this is a forecast, not a hindcast
When a label table carries an eval_ts, each user's history is truncated at that point
before embedding — for both the probe and the baseline — so metrics never peek at the
outcome window.
Where to go next
How it works
The encoder stack, temporal encoding, and the objective — what each piece means and why.
Bring your own data
The parquet data contract and how to point pragmatiq at real records.
The same library surface (pragmatiq/api.py) backs the CLI, notebooks, and production
callers, so anything quickstart does, you can drive stage by stage.