Use Julia, Python, and R for what they’re really good at — whatever that is for you. T orchestrates them.
Simulations in Julia, ML in Python, statistics in R — or the exact opposite. It doesn’t matter how you divide the labor: the hard part of polyglot data science was never the languages, it was the fragile seam between them.
A language for the LLM era, T is designed to be piloted by both humans and AI models. It gives you one hermetic dependency graph where your tools communicate without glue and execute consistently through space and time: on your laptop today, on a cluster tomorrow, and five years from now without bitrot.
Status: Version 0.55.4 “L’Ultime combat”.
Install Nix
(installs Nix and configures the rstats-on-nix cache in one
step):
curl --proto '=https' --tlsv1.2 -sSf -L https://install.determinate.systems/nix | \
sh -s -- install --no-confirm --extra-conf "
trusted-users = root $USER
substituters = https://cache.nixos.org https://rstats-on-nix.cachix.org
trusted-public-keys = cache.nixos.org-1:6NCHdD59X431o0gWypbMrAURkbJ16ZPMQFGspcDShjY= rstats-on-nix.cachix.org-1:vdiiVgocg6WeJrODIqdprZRUrhi1JzhBnXv7aWI6+F0="Try T immediately in an ephemeral shell:
nix shell --accept-flake-config github:b-rodrigues/tlangScaffold a project and enter its pinned environment:
t init --project my_project && cd my_project && nix develop(See the Nix Installation Guide and Getting Started Tutorial for full platform instructions).
Run t demo right in your terminal to see pipeline
introspection, hermetic Nix builds, Arrow in-memory inspection, caching,
and first-class error handling in action:
A complete analysis that simulates non-linear data in Julia, fits a gradient-boosted regressor in Python, plots ground truth vs predictions in R, and compiles a Quarto report:
p = pipeline {
-- 1. Simulate non-linear DGP in Julia (seeded)
sim_data = jln(
command = <{
using Random, DataFrames
Random.seed!(42)
t = 1:500
shock = cumsum(randn(500))
DataFrame(time = t, shock = shock, signal = sin.(t ./ 20) .+ shock .* 0.2)
}>,
serializer = ^ipc
)
-- 2. Train non-linear model & predict in Python (scikit-learn)
predictions = pyn(
command = <{
from sklearn.ensemble import HistGradientBoostingRegressor
X = sim_data[['time', 'shock']]
y = sim_data['signal']
model = HistGradientBoostingRegressor(random_state=42).fit(X, y)
sim_data['pred'] = model.predict(X)
sim_data
}>,
deserializer = [sim_data: ^ipc],
serializer = ^ipc
)
-- 3. Publication figure in R (ggplot2)
plot = rn(
command = <{
library(ggplot2)
ggplot(predictions, aes(x = time)) +
geom_point(aes(y = signal), alpha = 0.3, color = "#7f8c8d") +
geom_line(aes(y = pred), color = "#e74c3c", linewidth = 1) +
labs(title = "Julia Simulation + Python ML Predictions", y = "Value") +
theme_minimal()
}>,
deserializer = [predictions: ^ipc]
)
-- 4. Render reproducible Quarto report
report = node(script = "src/report.qmd", runtime = Quarto)
}
build_pipeline(p)
ggplot
object directly; T’s runner automatically renders and caches the visual
artifact without ggsave(). DataFrames pass between nodes
via Apache Arrow IPC (^ipc) without read.csv()
or to_csv() glue.<{ ... }> blocks. Nodes
accept external script files directly
(jln(script = "sim.jl"),
pyn(script = "train.py"),
rn(script = "plot.R")). Your Julia, Python, and R scripts
remain ordinary standalone files that your team can run or reuse
anywhere with standard tooling.raise), R (stop()), Julia
(error()), or T (error())—it does not crash
the entire pipeline build. T captures the error at the sandbox boundary,
serializes a structured VError artifact, and allows
independent branches to complete. Downstream nodes can inspect the error
with read_node() or explain(), or recover
programmatically.src/report.qmd into an HTML or PDF report inside the Nix
sandbox, directly embedding upstream metrics and figures.Most modern quantitative projects in research, central banks, official statistics, and regulated industries are polyglot by necessity: - Julia is unmatched for raw numerical simulation, ODEs, and heavy optimization loops. - Python is the standard for modern machine learning and deep learning tooling. - R remains the gold standard for survey statistics, econometrics, and publication-ready reporting.
Connecting them today forces you to choose between three bad options:
| The Status Quo | The Failure Mode |
|---|---|
In-process FFI (reticulate,
PyCall, RCall) |
Shared memory between multiple runtimes with competing garbage collectors and conflicting OpenMP/BLAS threads causes unexplained segfaults. Upgrading one runtime breaks the other. |
| Ad-hoc Bash scripts & CSVs | No caching: tweaking a title in an R ggplot re-runs your 3-hour Julia simulation. Column types and missing values silently mutate during CSV export. |
| Chained Docker containers | Huge container images, slow local development, impossible for an analyst to inspect or debug interactively on a laptop. |
^ipc) and standard model
serialization (^onnx, ^pmml,
^csv). No custom serialization glue scripts.VError
artifacts while independent parallel branches continue
uninterrupted.| Feature | {targets} | {rixpress} | Snakemake | Docker (packaging only) | T |
|---|---|---|---|---|---|
| Interface & Engine | R package (_targets.R),
host environment |
R package API, Nix engine | Python / CLI DSL, Conda/host | Container image, Docker daemon | Dedicated pipeline language, Nix engine |
| Cross-language seam | R-native (polyglot is bolted on) | R-native (Python nodes via
rixpress helpers) |
Shell scripts & CLI wrappers | Manual entrypoints & volume mounts | Process-isolated IPC across R, Python, and Julia |
| Intermediate I/O | Automatic | Automatic | Manual file paths | Manual volumes & files | Automatic (zero-boilerplate boundary transfer) |
| Node caching | Content-addressed (R) | Content-addressed (R) | Timestamp / file hash | Docker build layer cache | Content-addressed (all nodes) |
| System library locking | ❌ (Delegates to host) | ✅ (Hermetic Nix) | ⚠️ (Optional Conda) | ✅ (Per image) | ✅ (Hermetic per-node Nix sandbox) |
| Interactive inspection | ✅ (tar_read()) |
✅ (read_node()) |
⚠️ (File inspect only) | ❌ (Container attach) | ✅ (read_node(),
explain()) |
| Error resilience | ❌ (Aborts run) | ❌ (Aborts run) | ❌ (Aborts run) | ❌ (Container exits) | ✅ (First-class polyglot soft-failures) |
When you define a node using node(), rn()
(R), pyn() (Python), jln() (Julia), or
shn() (Shell), T treats the result as a first-class
Node object. These objects transition through two main
states:
build_pipeline(),
the node points to a concrete, immutable artifact in the Nix store.When you call read_node(p.node_name) in the REPL, T
looks at the node’s serializer and attempts to
automatically load the data back into the T environment:
| Serializer | Resulting T Type | Backend |
|---|---|---|
default /
serialize |
Varies | Native T binary serialization |
arrow |
DataFrame |
Apache Arrow IPC (zero-copy) |
csv |
DataFrame |
Native CSV parser |
json |
Dict / List |
JSON parser |
pmml |
Model |
Native T model evaluator |
explain()You can use explain() to look inside a built node:
-- Example: Inspecting a built R node
> model_node = p.model_r
> explain(model_node)
{
`kind`: "computed_node",
`name`: "model_r",
`runtime`: "R",
`path`: "/nix/store/...-model_r/artifact",
`serializer`: "pmml",
`class`: "lm",
`dependencies`: ["data"]
}
The path field is the escape hatch: it gives you the
absolute path to the node’s output in the Nix store. You can inspect the
artifact directly or pass it to external tools.
t update, and build a
hello-world pipelinefct_* helpers|>
forwarding semantics and short-circuiting^serializer system for data interchange and
materialization% shortcuts (%cd, %env, and
more)