Testing without a workspace¶
stevin generates SQL for a system most contributors can't run on a laptop, and that CI shouldn't need credentials to check. The suite is built in four layers, and each one is honest about what it can and cannot prove.
uv run pytest tests/unit # layers 1 and 2, and 3 once it is recorded — offline, every PR
uv run pytest -m integration # layer 4 — real workspace, nightly
| Layer | Proves |
|---|---|
| Unit tests | the pure middle of the pipeline does what it says |
| The fake warehouse | stevin's SQL matches stevin's reading of the manual |
| Transcripts | what Databricks answered, on the day they were recorded — built, and empty until a live run records them |
| The live suite | that it still does |
1. Unit tests¶
The type parser, loader, differ and planner are pure functions, so they are tested the
ordinary way: inputs in, values out, with golden plans in tests/snapshots/ for
anything shaped like a document.
This is where most of the tests live, and it is only possible because the middle of the pipeline does no I/O.
2. Convergence against a fake warehouse¶
tests/fake_warehouse.py is an in-memory Unity Catalog. It holds Table models,
answers the introspector's information_schema and DESCRIBE queries by rendering
them into the row shapes the real API returns, and interprets the statements the
planner generates by mutating those models.
That closes the loop offline:
plan = build_plan(diff(desired, live)) # what we would do
run(plan, fake) # do it
assert diff(desired, live_now(fake)) == () # nothing left to do
tests/unit/test_convergence.py makes that assertion for every kind of change —
creates, nested adds, renames, widenings, drops, constraints, reordering — and
tests/unit/test_executor.py uses the same fake to test skip, resume, failure,
locking and the destructive gate.
What this proves. That every statement stevin emits says what the change it came
from meant; that a plan closes the diff it was built from; that is_applied — the
executor's idempotency check — agrees with the differ; and that the executor's own
machinery behaves. In milliseconds, with no credentials.
What it does not prove
The fake implements stevin's reading of the Databricks manual. If we have misread it, the fake misreads it the same way and the tests still pass. It cannot tell you that Databricks accepts a statement, that a widening is permitted, or that a nested rename works.
Anything the fake doesn't recognise raises FakeSqlError rather than passing
quietly, so a new statement shape can't slip through untested — but a wrong
statement the fake also gets wrong will sail through. That is what layer 3 is for.
3. Transcripts: what the workspace actually answered¶
The layer above proves that stevin's SQL matches stevin's reading of the manual. Where that reading is wrong, the fake is wrong in the same direction and the offline suite agrees with the mistake. A transcript is the answer to that.
Built, and empty
No transcript has been recorded yet: tests/transcripts/ holds its README and
nothing else. The recorder and the replay are there and tested, and the test that
replays the recordings skips itself while there are none — so today this layer
proves nothing about Databricks. What follows is how it works once a live run has
written them.
One live run writes down every statement it sent and the rows or the error that came
back, per assumption. tests/unit/test_transcripts.py then replays each recording
through the probe it was recorded for — offline, with no credentials — so an assumption
keeps being checked against answers Databricks really gave, between live runs.
STEVIN_RECORD=tests/transcripts \
uv run pytest -m integration tests/integration/test_live_assumptions.py
Only a probe that held is written: one that didn't is something to fix in the code, and a transcript of it would assert the mistake. What is masked is only what is new on every run — the schemas the run makes, and the principal a grant names — and nothing else is touched, so the diff of a transcript is the diff of what Databricks answers. Read it.
A statement the transcript hasn't got fails, with the statement in the message: a
recording that no longer covers what stevin sends has stopped being evidence about
it, and needs recording again. See tests/transcripts/README.md.
A transcript says what was true when it was recorded. That is one thing more than the fake can say, and one less than the live suite — which is why the live suite stays.
4. Integration tests¶
tests/integration/ is the only source of truth about Databricks. Each test creates an
ephemeral schema, does its work, and drops it. They are marked @pytest.mark.integration
and skip themselves without credentials, so forks and laptops stay green, and they run
nightly in CI.
They assert the things nothing else can:
- a table created from a spec reads back as that spec;
- plan → apply → re-plan is empty, against the real thing;
DROP COLUMNreally is refused until column mapping is on;- every widening
widens()claims is one Databricks accepts; - the history and lock SQL —
MERGE,INTERVAL,current_user()— is valid; - dogfooding:
tests/messy_schema.pybuilds a schema the way a real team ends up with one — years ofALTERs, masks, a Python UDF, legacy partitioning, awkward names — andtest_live_dogfood.pyimports it, requires a plan of nothing but ownership claims, adopts it, and changes it. When you meet a real-world table stevin misreads, add its shape to the messy schema.
Every Databricks behaviour stevin relies on should have a test here and a link to the
documentation in its docstring. Where a behaviour is assumed but unverified, the code
says TODO(verify) rather than pretending.
What the live suite has not settled¶
Most of what stevin assumes about Databricks has been run against a workspace: the
TODO(verify) list was settled on 2026-09-19. These are still open, and each says so in
the source:
- A seed's load.
INSERT OVERWRITE … (columns) VALUES …is the documented grammar and a Databricks parser reads it, but no workspace has taken one from stevin (planner._load_seed).tests/integration/test_live_seeds.pyand the probe a seed's INSERT OVERWRITE with a column list is accepted settle it the next time the live suite runs. CLUSTER BY AUTOwithout predictive optimization. It was on in the workspace this was tested in (planner._clustering_clause); the probe CLUSTER BY AUTO is accepted and reads back answers it for the workspace in front of you.- Sending a step without waiting for it — not done yet.
applywaits up to 30 seconds for a statement's first answer, and an interrupt in that time has no statement id to cancel by;applysays so when it happens. Sending withwait_timeout="0s"is the Statement Execution API's documented way to get the id at once (introspect.WarehouseRunner). It changes how every step is sent, so it waits for a run against a workspace.
stevin verify runs the first two as probes in a workspace of your own.
The assumptions live in src/, not here¶
The behaviour assumptions themselves are stevin.probes.PROBES: a list of named
probes, each with its docs link, what stevin does because of it, and a check that
makes its own objects in a scratch schema. test_live_assumptions.py is a thin
parametrised caller of that list, and stevin verify runs the same list in a user's
workspace. So an assumption is written down once, and a user can settle it on the
runtime they actually have.
A new Databricks assumption therefore goes in probes.py, not in a test of its own —
unless what you are testing is stevin's own logic, which belongs in a test. The rule
of thumb: if the sentence is about what Databricks does, it is a probe; if it is about
what stevin plans, it is a test.
Offline, tests/unit/test_probes.py runs the probe machinery against the fake
warehouse. It deliberately asserts nothing about whether a probe holds: the fake
interprets stevin's own SQL, so a ✓ from it would be stevin agreeing with itself.
Running the live suite¶
Nothing in stevin has been verified against a real workspace until this has run. Every
TODO(verify) in the source names an assumption one of these tests settles.
You need:
- a catalog you can write to — every test creates a schema called
stevin_it_<random>in it and drops it, with everything inside, when it finishes; - a SQL warehouse — the tests run a few dozen small statements; a 2X-Small serverless warehouse is plenty;
- a principal allowed to
CREATE SCHEMAin that catalog andCREATE FUNCTIONin its schemas (the mask test creates a masking function); - optionally a principal to grant to —
account usersby default.
databricks auth login --host https://<workspace> --profile stevin-test
DATABRICKS_CONFIG_PROFILE=stevin-test \
DATABRICKS_WAREHOUSE_ID=<warehouse id> \
STEVIN_TEST_CATALOG=<scratch catalog> \
STEVIN_TEST_PRINCIPAL="account users" \
uv run pytest -m integration -v
A failure here is the point of the suite: an assumption about Databricks was wrong. Fix
the code, keep the test, and remove the TODO(verify) it settled — and if the fake
warehouse agreed with the wrong assumption, fix the fake too, so the offline suite stops
agreeing with it.
Asking one question of the workspace¶
A single live question doesn't need the whole forty-minute suite. Push a branch with the test on it and run the suite against it, filtered:
That's the honest way to settle a Databricks behaviour before building on it: probe it, read the answer, then write the code and the test that keeps it.
When the metastore says it is full¶
QUOTA_EXCEEDED.UC_RESOURCE_QUOTA_EXCEEDED — "Cannot create 1 Table(s) ... (estimated
count: 523, limit: 500)" — usually isn't. Unity Catalog
counts tables as they are created
and catches up with deletions later, so a day of live runs leaves the count far above
what is really there.
Two things keep it from happening, and one says so when it does:
- A test schema keeps nothing it drops.
ALTER SCHEMA … SET RETAIN DROPPED TO 0 HOURSturns off the seven-day recovery period, because a dropped table still counts against the quota for as long asUNDROPcould bring it back. Without this, a suite that makes a hundred tables a run fills a 500-table metastore in a couple of days while holding almost nothing. - Schemas a cancelled run left behind are swept. A new push cancels a running suite
mid-test, and its
stevin_it_*schema is never dropped; anything older than half an hour goes. - Then the quota is read, which is what asks Unity Catalog to recount it. If it is still at the limit the suite skips rather than failing every test in it — with one line saying why.
If the count stays above what the catalogs really hold, it is dropped tables still inside their recovery period, from before the setting above. They age out; a metastore that needs to run this suite often is worth asking your Databricks account team to raise the table quota on.
The docs' pictures¶
Every terminal on the docs site is stevin's own output. tests/screens.py writes a
small project for each scene, applies its "before" specs through the CLI into the fake
warehouse, runs the command the page shows, and records what the CLI printed as an SVG.
The spec files a page quotes are written alongside, so the YAML next to a plan is the
YAML that made it.
test_screens.py fails when a picture no longer matches what the CLI prints. After
changing anything a user sees, remake them and look at the diff:
A new feature earns a scene in the feature gallery: add it to
tests/screens.py and a section to docs/features.md.
Where a new test goes¶
| You changed | Test it in |
|---|---|
| The type parser, loader, differ, planner | tests/unit/, with a snapshot if it shapes a plan |
| The SQL a step generates | tests/unit/test_convergence.py — teach the fake the statement |
| The executor, history, locking | tests/unit/test_executor.py with MemoryHistory |
| An assumption about what Databricks does | a probe in src/stevin/probes.py, with the docs link — then record it |
| Anything a user sees in the terminal | a scene in tests/screens.py, then mise run screens |