pragmatiq
Get started

Quickstart

Install pragmatiq and run the full pipeline — generate data, tokenize, pretrain, embed, and probe — on a GPU or a laptop CPU.

pragmatiq is an independent implementation inspired by the PRAGMA paper (arXiv 2604.08649) and is not affiliated with or endorsed by Revolut.

This runs the entire pragmatiq pipeline end to end on synthetic data. By the end you will have a trained model and a credit-risk probe score measured against a raw-count baseline — proof the embedding carries signal beyond trivial counts.

Requirements

Python 3.11+, torch 2.6+. A GPU is picked up automatically when visible (see Running on GPU); on a CPU the fast path below takes a few minutes in fp32. The commands use a virtual environment so the install is isolated.

Install

git clone https://github.com/dynamiq-ai/pragmatiq.git
cd pragmatiq
python3.11 -m venv .venv && source .venv/bin/activate
pip install -e ".[train]"

The train extra brings in Lightning, which the pretraining stage runs on; the plain install is the slim inference core (validate, tokenize, embed, probe with a trained run). Other extras add focused tooling — .[serve] for ONNX/Triton export, .[aml] for the transfer-graph GNN, .[text] for the frozen text embedder, .[tracking] for experiment tracking, .[demo] for the Streamlit app, .[full] for all of them — see Install for the full table.

Run the whole pipeline

pragmatiq quickstart --n-users 2000 --max-steps 80   # the fast path: a few minutes
pragmatiq quickstart                                 # the reference run (50k users, 400 steps)
from pragmatiq import api

result = api.quickstart(n_users=2000, max_steps=80)
print(result["message"])

quickstart runs five stages in order:

  1. generate synthetic users and event histories,
  2. fit the key–value–time tokenizer,
  3. pretrain a nano masked-language model,
  4. embed users,
  5. probe a gradient-boosting credit-risk classifier against a raw-count baseline.

Read the result

The run prints a one-line summary like:

credit probe beats the raw-count baseline  ·  run: runs/quickstart

The probe head is gradient boosting by default (HistGradientBoostingClassifier), and the raw-count baseline uses the same classifier — so the gap reflects the representation, not the model family. Both ROC-AUC and PR-AUC are reported (PR-AUC is the honest headline on low-prevalence risk tasks).

Why this is a forecast, not a hindcast

When a label table carries an eval_ts, each user's history is truncated at that point before embedding — for both the probe and the baseline — so metrics never peek at the outcome window.

Where to go next

The same library surface (pragmatiq/api.py) backs the CLI, notebooks, and production callers, so anything quickstart does, you can drive stage by stage.

On this page