Running tests inside a session
An agent session that changes behaviour should be able to run the test that covers it,
before it pushes. For a long time it could not: bin/rails did not boot at all, so every
form of verification was deferred to CI, and the first CI run was where a session found
out whether the test it had just written even ran
(#592).
It works now. This page is the whole recipe, and the reasoning for the two pieces of it that are not obvious.
The gems are already there
Section titled “The gems are already there”Nothing to install, and nothing to export. A clone of this repo at a commit that has not
touched the Gemfile shares the image’s /usr/local/bundle, and AgentSessionJob settles
that inline while it sets the clone up — so .bundle/config is on disk before the agent
process exists:
bin/rails --version # => Rails 8.1.3If that instead prints Bundler::GemNotFound listing a hundred gems, the clone needs a
bundle of its own (its Gemfile differs from the image’s) and BundleInstallJob is
installing one in the background. It takes a couple of minutes; the session log says when
it is done. bin/agent-dev --skip-server forces the same repair in the foreground.
Point at the dev database
Section titled “Point at the dev database”There is no Postgres on localhost and no way for a session to start one — no root, no
sudo, no server binaries in the image. What there is, is the shared devdb Kamal
accessory on the container network, which is the same database
bin/agent-dev boots against. Borrow its wiring:
export DATABASE_HOST=zimmer-devdb \ DATABASE_PORT=5432 \ DATABASE_USERNAME=zimmerdev \ DATABASE_PASSWORD=zimmerdev \ DATABASE_SSLMODE=disable \ DATABASE_NAME="$(bin/agent-dev --print-database-name)" \ RAILS_ENV=testunset DATABASE_URLThree of those lines are load-bearing and none of them is guesswork:
DATABASE_SSLMODE=disable. The deployment setsrequirefor DigitalOcean Managed Postgres;devdbspeaks plaintext, and libpq’s answer to that pairing is a refused connection that reads like a broken database.DATABASE_NAMEfrombin/agent-dev --print-database-name. Every session on the droplet shares onedevdb, so the name is derived from the clone directory. Skip it and two sessions run each other’s migrations.unset DATABASE_URL. It outranks every discrete variable above, and Rails’ owneach_current_configurationdrops the test databases whenever it is set.
Then create the schema once per clone:
bin/rails db:test:prepare # ~25sRun the test
Section titled “Run the test”bin/rails test test/jobs/bundle_install_job_test.rbTargeted, by file or by file:line. Sessions run with PARALLEL_WORKERS=2, which Zimmer
sets, and a session also runs inside a 4 GiB memory cgroup of its own — so a run that dies
with exit 137 was killed by that bound, and the answer is to run one file at a time
rather than to raise the worker count.
Do not read the per-session bound as protection for anything but this session. Three sessions each running a parallel suite, none of them near its own 4 GiB, once used 8.7 GB between them and killed the Rails worker; what protects the worker and the other sessions is the 6 GiB cap on the sessions pool as a whole, not the per-session one. See All sessions together get a second bound.
The full suite is not the point of this path and is not worth attempting — it is thousands of tests against a shared database on a busy droplet. CI runs it. Run what your diff touches, and let CI be the one that runs everything.
System tests work too
Section titled “System tests work too”Chrome is in the image. It cannot start its setuid sandbox in this container, and
test/application_system_test_case.rb already carries the opt-in for exactly that case —
CHROME_NO_SANDBOX=true, which is deliberately not spelled CI=true (that would also
point the binary at a chromium-browser path only the CI image has).
One extra step first: the layout links a compiled stylesheet, and a fresh clone has none,
so every page render fails with The asset 'tailwind.css' was not found in the load path.
bin/rails tailwindcss:build # once per cloneCHROME_NO_SANDBOX=true bin/rails test test/system/code_block_copy_test.rbScreenshots land in tmp/capybara/ — Capybara writes one automatically for every failing
system test, which is usually the fastest way to see what a session’s change actually did.
Share them the way references/GIT_WORKFLOW.md
describes: a session’s filesystem is not reachable by the person reading the PR.
A system test that drives a runtime rather than a screen is a different matter —
test/system/smoke_test.rb spawns a real agent process and will not pass here. That is not
a broken recipe; it is a test that needs something this environment does not have.
Put anything long-lived in the scratch directory
Section titled “Put anything long-lived in the scratch directory”The session container is recreated on every deploy, and the clone goes with it. Anything
that has to outlive that belongs under $AO_SESSION_SCRATCH_DIR, which is on the durable
volume. Test databases do not need this — they live in devdb — but a fixture corpus or a
captured artifact does.
What this does not cover
Section titled “What this does not cover”- Booting the app to click through a change: that is
bin/agent-dev, which prepares the same database in the development environment and starts a server. - The full suite,
bin/rubocop --parallelover the whole tree, or Brakeman. CI owns those, and it is faster at them.