Nightshift: Coding while you sleep

I have written about my nightshift with claude code before.
TLDR
I’m working heavily with Matt Pococks skills. Did my own extensions. My workflow does now 20-30 tickets a night while I’m sleeping. It’s a process to get there. You can start as well.
Move all the thinking to the front
Very much different workstyle for all developers.
Yellow is me working with Claude. Blue is Claude alone. I spend 2-3 hours reviewing, refining, grilling and triaging tickets. Claude keeps working for 8-14h.
Getting ready for the nightshift
In the very beginning I start with an idea. It goes through /grill-me and /to-issues. It records decisions, creates a GLOSSARY.md and all nessesary context.
The questions it asks me during /grill-me are sometimes painful and hard to answer. Do not skip or just accept the defaults (although there are often very good). You can always chat
about them and this helps you to understand the question better.
What I like about /to-issues that it cuts the problem vertical, not horizontal. Each ticket is a tracer bullet: a narrow path through every layer, from database schema to API to UI to tests.
It also tries to keep the ticket concerned with one issue to fit well in a context window. A ticket describes behaviour and acceptance criteria. It avoids file paths and code snippets, because those go stale while the ticket waits in the queue.
Once this is done, all tickets get the ready-for-agent label. The nightshift script can pick those up and will implement it.
Before I start, I checkin with Claude to make sure the tickets have the right dependencies the right build in order. Every ticket system has a mechansim for that. Sometimes I set ‘cluster’ labels to a bunch of tickets if I want to have those explicitly build and not all ‘ready-for-agent’ tickets.
Everything runs on Opus 5.5 so far. I have chatted with Claude about using Sonnet, but the arguments pros / cons from Claude itself let me stick with Opus.
Nightshift, the overnight loop
Nightshift is one shell script. So far I have no one fits all solution, it is always modified for each project.
The script runs until the ticket system has no ready-for-agent tickets anymore.
Examples below are from my current projects writing a server following a public specification in Rust with a Postgres database.
Preflight
Before the first ticket, the script checks that the night can succeed. It verifies the required tools, confirms the tracker is reachable, refuses to start on a dirty working tree and switches to a dedicated nightshift branch. It starts a throwaway Postgres in Docker for the integration tests. It writes the sandbox configuration and runs a smoke test inside it: can the sandbox build Rust, reach the package registry and connect to the database? A failure here stops the run before it wastes a single ticket.
The sandbox
Every agent runs headless (claude -p), in the claude sandbox runtime, with permission prompts disabled. Disabling prompts is safe only because the sandbox is tight:
- Network: only the model API and the package registry. No access to the issue tracker.
- Filesystem: writes only to the project and the build caches. SSH keys, GPG keys, Docker and the tracker credentials cannot be read.
- The spec is read-only. The protocol specification lives in a submodule the agent cannot change. An agent cannot “fix” a failing test by editing the standard. This is my current case.
The agents never touch the tracker. The script fetches the ticket text into a work folder before the agent starts, and does every comment, label change and close itself afterwards. The agent needs no tracker credentials at all.
Picking the next ticket
Each round of the loop picks one ticket.
Dependencies always beat priority: a high-priority ticket waits behind a low-priority ticket it depends on.
Because the script re-reads the tracker every round, the queue is ‘alive’. This means even after a nightshift started I can still add ready-for-agents tickets to the queue and they will get done.
Working one ticket: three fresh agents
Each ticket passes through up to three agent sessions. Each session starts with a fresh context. No agent sees another agent’s conversation, only its commits and notes.
All three prompts share a common context section.
This informs the agent that no human is available, so it must never ask a question. If anything is unclear, the agent selects the option that most closely matches the specification and the ADRs, and records this choice in its summary. It lists what must be read before every change: the project rules, the coding standards, the glossary, the relevant ADRs, and the cited sections of the specification. It describes the sandbox so that the agent does not waste time starting Docker. It lists the quality checks that must pass. And it defines the termination condition: The last line of the output must read either “done” or “needs info,” followed by a one-line explanation.
Context for all following agents
# Nightshift context (shared by every phase)
You are an autonomous agent working overnight on **t-project**, a reference T-Project written in Rust. No human is available, so never ask questions and never wait for input. When something is ambiguous, choose the option most faithful to the T-Project SAI spec and the ADRs, and record the choice in your summary.
## Ticket
- Ticket: **#{{TICKET}}**, full text including comments in `{{WORK}}/issue.md`.
- Parent PRD: `{{WORK}}/prd.md`.
- Fixed point (the commit before this ticket's work): `{{BASE}}`.
- Branch: `{{BRANCH}}`. Commit locally only. Never push, never rewrite history at or before `{{BASE}}`, never switch branches.
## Read before you change anything
1. `CLAUDE.md`, `docs/CODING_STANDARDS.md`, `GLOSSARY.md`.
2. The ADRs in `docs/adr/` that the ticket or the PRD mentions.
3. The t-Project spec sections the ticket cites, in `t-project/architecture/` (`§2.x` = `03-credential-format.md`, `§3.x` = `04-verification.md`, `§7.x` = `07-t-authority-apis.md`), plus the OpenAPI, schemas, type metadata and test vectors in that directory. `t-project/` is read-only.
4. The existing code in the areas you will touch. Follow its patterns.
## Environment (sandboxed)
- No Docker, no Gitea access, no `sudo`. Network access is limited to crates.io, GitHub and Anthropic.
- Postgres: use `DATABASE_URL` (already set; it is a Unix-socket connection). Use `#[sqlx::test]` for per-test databases. Use `cargo sqlx prepare --workspace` to refresh `.sqlx/`, installing `sqlx-cli` with `cargo install sqlx-cli --no-default-features --features postgres,rustls` if it is missing.
- SoftHSM2: the module is at `SOFTHSM2_MODULE`. Tests create a temporary token directory with their own `SOFTHSM2_CONF`.
- Binding to `127.0.0.1` works for local test servers (DNS and HTTPS stubs).
- Some things can only be verified in CI or with Docker: the image build, `docker compose`, the Gitea Actions workflow, and the end-to-end job. Write them carefully, validate what you can statically (for example `docker compose config` is unavailable, so re-read the YAML), and list them under "Unverified" in your summary.
- Gates. These must pass before you finish:
{{GATES}}
## Outputs the nightshift script reads
- `{{WORK}}/summary.md`: Markdown for the issue comment. Include: what was built, the module interfaces and seams tested, spec or ADR deviations (also recorded in `CONFORMANCE.md`), decisions you took on ambiguity, what is unverified, and follow-ups.
- `{{WORK}}/followups.jsonl` (optional): one JSON object per line, `{"title": "...", "body": "...", "labels": ["needs-triage"]}`, for issues the script should create. Use it for real follow-up work, not for things you should have done.
- Your **final line** must be exactly one of:
- `<result>DONE</result>`
- `<result>NEEDS_INFO: <one-line reason></result>`: the ticket cannot be implemented as written (contradiction, missing decision, impossible prerequisite). In that case revert the working tree to `{{BASE}}` with `git reset --hard {{BASE}}`, and explain in `summary.md` exactly what decision is needed.
The implement agent works in four moves:
- Plan. It writes a plan file: the acceptance criteria as a checklist, the spec requirements involved, the modules it will touch and the seams it will test at. It uses a
mattpocock-skills:codebase-designskill to choose deep modules. In an interactive session, I would approve the seams. At night, the ticket’s acceptance criteria are the pre-agreed seams. - Build test-first. It follows
mattpocock-skills:tddskill, one red-to-green slice at a time. Expected values come from the spec and its test vectors, not from the code under test. - Complete. It updates the conformance tracker (my own tracker, a markdown file, to make sure I follow the public spec), the API description and the glossary as needed. If needed a new ADR is written.
- Commit with a summary and any follow-up tickets it wants filed (we come to that later)
Plan & Build Prompt
# Phase: IMPLEMENT ticket #{{TICKET}}
{{CONTEXT}}
## Process
This follows the `implement` skill (which cannot be invoked headless), with the TDD and design skills invoked for real.
1. **Understand.** Read everything listed above. Write `{{WORK}}/plan.md` with:
- the acceptance criteria restated as a checklist
- the spec requirements involved (with § references)
- the modules you will add or change and their interfaces
- the **seams you will test at**
Invoke the Skill tool with `mattpocock-skills:codebase-design` and use its vocabulary to choose deep modules and place the seams. There is no user to confirm the seams with: the ticket's acceptance criteria and this plan are the pre-agreed seams. Write them down before writing any test.
2. **Build test-first.** Invoke the Skill tool with `mattpocock-skills:tdd`, and follow it at the seams from your plan, one vertical red → green slice at a time. Expected values come from the spec, the test vectors and worked examples. Run `cargo check` and the single test you are working on often.
3. **Complete the ticket.** Every acceptance criterion is met, or explicitly marked as deviating with a reason. Update `CONFORMANCE.md`, `openapi/management-api.yaml` (for management or OAuth endpoints), and `GLOSSARY.md` if you introduce a domain term. If you make a decision an ADR does not cover, invoke the Skill tool with `mattpocock-skills:domain-modeling` and record a new ADR.
4. **Run the full gates** listed above. Fix until they are green. Never weaken a test or lint to pass.
5. **Commit.** Use one or more Conventional Commits like `feat(#{{TICKET}}): ...`. The body includes `Refs #{{TICKET}}` and ends with:
`Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>`
Leave the working tree clean. Do not commit anything under `.nightshift/`.
6. **Write `{{WORK}}/summary.md`** and any `followups.jsonl`, then print the result line.
Scope discipline: implement **only** ticket #{{TICKET}}. If a later ticket covers something, leave a clean seam or a stub for it, as the ticket text says. Do not build ahead.
The prompt also enforces scope: build this ticket only. If a later ticket covers something, leave a clean seam for it.
The review agent comes in with fresh eyes and is told to be demanding. It runs mattpocock-skills:code-review skill that spawns three parallel reviewers:
- Spec: does the code do what the ticket, the PRD and the cited spec sections require?
- Standards: does it follow the coding standards and the glossary?
- Security and correctness: authorisation gaps, data leaking between tenants, race conditions, crypto misuse, secrets in logs, missing audit events.
It writes every finding down, fixes all spec and security findings and every hard standards violation, and records why it left anything unfixed. A bug fix starts with a failing test.
Review Prompt
# Phase: REVIEW and FIX ticket #{{TICKET}}
{{CONTEXT}}
Another agent implemented ticket #{{TICKET}} in the commits `{{BASE}}..HEAD`. Its notes are in `{{WORK}}/plan.md` and `{{WORK}}/summary.md`. You are the reviewer, with fresh eyes. Be demanding: this code is a reference implementation that other people will learn from.
## Process
1. Invoke the Skill tool with `mattpocock-skills:code-review`, using this input:
- **Fixed point:** `{{BASE}}`.
- **Spec source:** `{{WORK}}/issue.md` plus `{{WORK}}/prd.md`. Tell the Spec sub-agent to also check the implementation against the T-project spec sections the ticket cites (`t-project/architecture/`) and the relevant ADRs.
- **Standards sources:** `docs/CODING_STANDARDS.md`, `CLAUDE.md`, `GLOSSARY.md`.
- In addition to the two axes, have a third parallel sub-agent do a **security and correctness** pass. It looks for authentication and authorisation gaps, cross-Operator data leaks, TOCTOU and race conditions (spec-mandated uniqueness, single-use challenges), crypto misuse (ES256 R‖S encoding, JWK thumbprints, constant-time comparison), SSRF, secrets in logs, and missing audit events.
2. Write all findings to `{{WORK}}/review.md`, grouped by axis.
3. **Fix** every Spec finding, every hard Standards violation, and every security finding. Fix judgement-call smells when the fix clearly improves the code. For each finding you decide not to fix, add a line in `review.md` saying why. For refactors, invoke the Skill tool with `mattpocock-skills:codebase-design`. Keep tests at the seams. Fixing a bug starts with a failing test that reproduces it.
4. Run the full gates until they are green.
5. Commit the fixes as `refactor(#{{TICKET}}): address review findings` or `fix(#{{TICKET}}): ...`, with `Refs #{{TICKET}}` and the Co-Authored-By trailer:
`Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>`
Leave the tree clean. If nothing needed fixing, make no commit.
6. Append a "Review" section to `{{WORK}}/summary.md`: the number of findings per axis, what you fixed, what you consciously left and why. Then print the result line. Use `NEEDS_INFO` only if the review shows that the ticket itself is contradictory. In that case do **not** reset; explain in `summary.md`.
The gates run outside any agent: formatting, the linter with warnings as errors, and the full test suite. The script runs them itself, this is deterministic, no need for an agent.
The repair agent runs only if the gates are red. It gets the gate log and one instruction: find the root cause. No skipped tests, no silenced lints. If it thinks a test is wrong, it must cite the spec line that proves it. It gets one attempt.
Repair Prompt
# Phase: REPAIR failing gates for ticket #{{TICKET}}
{{CONTEXT}}
The nightshift script re-ran the gates after review, and they failed. The output is in `{{WORK}}/gates.log`. Find the root cause and fix it properly: no skipped tests, no `#[ignore]`, no `allow` lints without a justification. If a test is wrong rather than the code, cite the spec line that proves it. Run the gates until they are green, then commit as `fix(#{{TICKET}}): ...` with `Refs #{{TICKET}}` and the Co-Authored-By trailer:
`Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>`
Append a "Repair" note to `{{WORK}}/summary.md`, then print the result line.
Between phases, the script runs an integrity check. Did the agent rewrite history? Switch branches? Touch the spec? Leave uncommitted files? Any of these fails the ticket.
When things go wrong
- Needs info. The agent found a contradiction or a missing decision. Its work is discarded, the open question goes on the ticket as a comment, and the ticket leaves the queue. Its dependents stay blocked. This is the system working as intended: a question asked is cheaper than a guess built on.
- Failed. The agent timed out, the gates stayed red after repair, or an integrity check failed. The partial work is kept on its own branch for a human to inspect, and the main nightshift branch goes back to where it was before the ticket.
- Usage limit. The model provider’s usage window ran out. That is not the ticket’s fault, so the tracker is left alone. The script parks the partial work, sleeps until the limit resets and starts the ticket again from a clean state. A ticket that hits the limit three times is too big for one window and goes to a human with a suggestion to split it. I ran into that and had to add this.
The night stops on its own in four cases:
- No eligible tickets left
- An optional ticket cap reached, I can also run
./nightshift.sh 1and it will only do one ticket. - A login failure or a usage limit that won’t reset in time
- Two hard failures in a row. Two consecutive failures usually mean something systemic is broken, such as the sandbox or the database, and burning through the rest of the queue would only mislabel good tickets.
I run everyting in a dedicated Linux VM. The script runs in tmux. At some point this will evolve into an agent itself. I’m getting closer.
What does it look like
[05:17:03] === ticket#108 (18) ===
[05:17:03] implement: agent running (timeout 3h, transcript /home/oliver/Source/project-t/.nightshift/tickets/108/implement.jsonl)
[05:37:45] review: agent running (timeout 3h, transcript /home/oliver/Source/project-t/.nightshift/tickets/108/review.jsonl)
[05:56:15] gates: running in sandbox
[05:59:53] follow-up filed: http://git.internal/oliver/project-t/issues/177
[05:59:53] #108 closed
[05:59:53] === ticket#119 (19) ===
[05:59:53] implement: agent running (timeout 3h, transcript /home/oliver/Source/project-t/.nightshift/tickets/119/implement.jsonl)
[06:12:00] review: agent running (timeout 3h, transcript /home/oliver/Source/project-t/.nightshift/tickets/119/review.jsonl)
[06:23:56] gates: running in sandbox
[06:27:26] #119 closed
[06:27:27] === ticket #120 (20) ===
[06:27:27] implement: agent running (timeout 3h, transcript /home/oliver/Source/project-t/.nightshift/tickets/120/implement.jsonl)
[06:45:25] review: agent running (timeout 3h, transcript /home/oliver/Source/project-t/.nightshift/tickets/120/review.jsonl)
[07:00:17] gates: running in sandbox
[07:00:21] gates red; one repair session
[07:00:21] repair: agent running (timeout 3h, transcript /home/oliver/Source/project-t/.nightshift/tickets/120/repair.jsonl)
[07:04:41] gates: running in sandbox
[07:08:05] #120 closed
[07:08:06] === ticket #121 (21) ===
[07:08:06] implement: agent running (timeout 3h, transcript /home/oliver/Source/project-t/.nightshift/tickets/121/implement.jsonl)
[07:24:25] review: agent running (timeout 3h, transcript /home/oliver/Source/project-t/.nightshift/tickets/121/review.jsonl)
[07:46:07] gates: running in sandbox
[07:49:40] #121 closed
The morning
Inn the morning I look at the tickets. It adds comments and labels as needed.
- Closed tickets have a comment with the agent’s summary: what was built, which seams were tested, which decisions it took on ambiguity, what it could not verify inside the sandbox and the list of commits.
- Needs-info tickets have a precise question.
- Ready-for-human tickets point at a branch with the partial work and the last gate output.
- New needs-triage tickets are follow-ups the agents filed for real work they found but were told not to build.
Nothing is pushed. All work sits on the nightshift branch until I have reviewed it and merged it. Since I work alone here I merged it to main. I could also have a PR here or a PR after each successfull ticket. Not 100% sure if this would be useful even in a team setup, since this is potentially a lot of code to review. In the current project I do 20ish tickets a night.
The triage tickets are sometimes very short and conise and I had a hard time understaning them. So I came up with just another skill which enhances the tickets with a few things:
- User impact: one concrete scenario from the user’s life (who, where, doing what), then what they experience as bullets. Say plainly whether data is lost.
- Business impact: cost with an honest order of magnitude, quality perception, cleanup or support burden. When the product’s documented plans (scale, hosting, audience) change the answer, give both today’s and that future’s answer.
- Consequence of not fixing: what keeps working, what happens and how often. When frequency is unmeasured, say so.
- Add a priority to the ticket: high, medium, low
This makes it so much easier to judge when things should be implemented. I do have now a lot of low prio tickets which I will not do, but all medium and higher get triaged and ready-for-agent.
Learnings
I’m quite happy with that so far. I know I will take it further but this is the next step
Specification quality is the bottleneck. The agents are good at building what a ticket says. They are bad at guessing what it means. Every hour in the grill and the triage pays back many times over. Yes, painful but worth it.
Fresh contexts beat long sessions. Separate implement, review and repair agents are it. They do not defend their earlier decisions. Smaller context.
The script decides, the agents build. Picking tickets, running the gates, judging integrity and writing to the tracker are deterministic.
Clusters make mornings reviewable. I introduced clusters to have a better context when reviewing tickets. It make the story better for me.
User/Business Impact, consequence of not fixing Gamechanger. Better understanding of all implication.
Does this work for you?
Maybe not, maybe yes. Your environment is for sure different.
I think everyone can start building a system to move it that direction. What are the borders for you what is possible in your setting? Find out!
I know there are still a few things missing and things I did not mention what I have, like my CI/CD chain.
I went from ChatGPT -> Code Completion with cursor -> Claude -> Spec driven interactivly -> A few tickets at once -> Many tickets over night.
This will take a while and each step will build trust and shows you where it works and where it breaks. Keep experimenting.
Enjoy the journey.