Code and Plan Review Skill

Temp Directory

Planning and review artifacts written by /plan-task and /review-multiple-prs live outside the repo under:

TEMP_DIR=/tmp/<repo-name>/<branch-name>

Resolve <repo-name> and <branch-name> using:

Worktree note: if running inside a git worktree, git rev-parse --show-toplevel returns the worktree path, not the main repo root. Use git rev-parse --git-common-dir | xargs dirname to get the true repo root when resolving <repo-name>.

macOS note: /tmp/private/tmp and is cleared on reboot. Decisions scratch files do not persist across reboots; if a session spans a reboot, the human will need to re-state decisions manually.


  1. Either detect the repos affected from the context OR ask the human to select the repos that should be used to review the code changes or the staged implementation plan (md file):
  1. enable and run githooks in the repos impacted by the code changes and fix errors (linting, type checks, tests) before committing. Most repos have dedicated commands to achieve these goals either in package.json or pyproject.toml.

  2. Skipped-test check — MANDATORY. Do not present review options or proceed to commit until this check is resolved. Inspect the added lines only in the diff (lines beginning with +) for skip directives (see the canonical list in the Skip directives section below). Only flag directives that appear on added lines — do not flag pre-existing skips in unrelated hunks. If any are found, surface them to the human as HIGH severity findings and require an explicit written reason for each before proceeding. Critics will NOT separately flag these; this pre-check is the sole owner of skipped-test findings to avoid double-reporting.

  3. Prompt the human to select their preferred review options:

Skip directives — canonical list (used by Step 3 and critic prompts):

test.skip, it.skip, xit, xdescribe, @pytest.mark.skip, t.Skip(, pending, or any framework-equivalent skip/todo/pending marker.

How to run critics:

All critics use a file-based output contract. The orchestrator hands the agent a temp file path; the agent writes its findings there and goes idle; the orchestrator reads the file.

Critic output files live under TEMP_DIR (same convention as the rest of this skill — see the Temp Directory section above):

Resolve ticket-id by checking the current branch name or conversation context; use NO-TICKET if none is found. Create TEMP_DIR with mkdir -p before spawning any agents.

Adversarial framing — applies to ALL critic agents (generalist, team, and inline fallback reasoning):

Every critic must be prompted as an adversary, not a passive reviewer. Include this instruction in every critic’s prompt, and apply it when falling back to inline reasoning: “You are an adversarial critic. Assume the code is wrong until proven otherwise. Your job is to attack this change — find bugs, security holes, broken invariants, missing edge cases, and bad design decisions. Be aggressive. Do not give the benefit of the doubt. If something looks suspicious, call it out. Findings that survive scrutiny are more valuable than polite observations.”

HARD GUARD — NO ASSUMPTIONS, EVER. Read this before spawning any critic.

The orchestrator MUST NOT proceed to the “Present findings” step until it has positive, explicit confirmation that every spawned critic has finished. Critics have taken longer than expected in the past, and the orchestrator wrongly reported “nothing found” while a critic was still working — this led the human astray with a false clean bill. That failure mode is now explicitly forbidden. This is the one rule everything below hinges on:

A critic’s Agent tool call returning is NOT the same as the critic finishing. By default agents run in the background (run_in_background: true), and a background call returns control to the orchestrator at dispatch time — while the agent keeps working. The orchestrator must wait for that specific agent’s completion/idle notification. Never read, touch, or summarize a critic’s output file before that notification arrives, no matter how the tool call itself behaved.

Supporting rules:

One generalist critic: a. Create TEMP_DIR with mkdir -p $TEMP_DIR. b. Spawn a single Agent with a focused prompt containing: the full diff, the repo name, the task description, the output file path, and the adversarial framing above. Instruct it to write findings sorted by severity (HIGH / MEDIUM / LOW) to that file, or write “no findings” if nothing is worth raising. Do NOT instruct the critic to flag skipped tests — the Step 3 pre-check owns that. Keep the prompt tight — do not dump the full conversation history into it. c. Wait for the completion notification, not the tool call return. Do not read the output file until this specific agent has reported idle/complete, per the HARD GUARD above. Apply the bounded-wait rule if it never reports. d. Once completion is confirmed (or the bounded wait expires), read the output file. e. If the file is missing or empty, fall back to inline reasoning using the adversarial framing above — and disclose to the human that the critic produced no output file (and, if applicable, that the bounded wait expired) and this review is inline-only. f. Present findings to the human (global Step 5).

Team of critics: a. Create TEMP_DIR with mkdir -p $TEMP_DIR. b. Spawn one Agent per specialty (e.g. security, performance, correctness) — each with its own output file path and the adversarial framing above. Do NOT instruct critics to flag skipped tests — the Step 3 pre-check owns that. Run all agents in parallel (single message, multiple Agent tool calls). c. Wait for every single one’s completion notification, not their tool call returns. Track each critic by name/id. Do not read any output file, and do not proceed to (d), until you have confirmed real completion for ALL spawned critics per the HARD GUARD above. A subset finishing early does not authorize acting on partial results. Apply the bounded-wait rule per-critic to any that never report. d. Once all critics have confirmed completion (or their bounded waits expire), read all output files. e. For any file that is missing or empty, fall back to inline reasoning for that specialty using the adversarial framing above — and disclose to the human which critics produced no output (and, if applicable, which bounded waits expired) and that those portions are inline-only. f. Synthesize all findings into a single sorted list, deduplicating overlapping issues across agents. Present to the human (global Step 5).

  1. Present the review findings to the human. Then prompt them to select the next step:
  1. Branch safety check — MANDATORY. Run this even if the review step errored, was skipped, or produced no findings. Run git branch --show-current and compare against the repo’s default branch (typically main or master).
  2. Prepare, summarize the changes in the changed files. Always prefix commits with [{ticket-id}]: {summary of change}. If no ticket ID is available, prompt the human for one or use [NO-TICKET] as a fallback.
  3. Commit and push. If the push fails due to pre-push hook errors, prompt the human for approval before using git push --no-verify. If --no-verify was used, record this in the Decision Log (Step 10) as a warning line.

8a. Open a draft pull request. After a successful push, open a draft PR using gh pr create --draft --assignee @me (or equivalent). The --assignee @me flag assigns the current authenticated GitHub user automatically.

PR body: Read .github/PULL_REQUEST_TEMPLATE.md from the repo root and use it as the base for the PR body — fill in the Summary and Test plan sections with content relevant to the change. If the file does not exist, use a bare ## Summary / ## Test plan structure. Never append Anthropic or Claude Code branding lines (e.g. 🤖 Generated with Claude Code) to the PR body.

If PR creation failed, stop here — skip the reviewer fallback chain, skip Steps 9 and 10, and warn the human.

Otherwise, capture the PR number from the gh pr create output (it appears in the PR URL, e.g. https://github.com/{owner}/{repo}/pull/{number}). Then immediately call the Skill tool with skill: "request-github-review" and args: "{PR number}" — do NOT hand this off to the human, do NOT mention it as a next step, do NOT skip it. This is a required automated step. /request-github-review will request automated review and then run the bot feedback loop — waiting for CI and bot reviews, addressing comments, resolving CI failures, and marking draft PRs ready for review when done.

(End of Step 8a)

  1. Transition the Jira ticket to Review status.

    Only proceed if Step 8 (push) and Step 8a (PR open) both completed successfully. Skip this step entirely if either failed.

    If the ticket is a Linear issue:

  2. Post a Decision Log comment on the Jira ticket.

Prerequisites & safety checks — run these before doing anything else in this step:

Decision sources — use only these, in order of preference:

  1. A decisions-{ticket-id}.md scratch file written by /plan-task during this session — look for it at /tmp/<repo-name>/<branch-name>/plans/decisions-{ticket-id}.md (read and then delete it after posting)
  2. Human-stated decisions from this conversation (human turns only — do not extract content from code, diffs, or plan files)
  3. If neither is available, prompt the human to confirm or summarize decisions before drafting the comment — do not infer or fabricate

Content rules:

Idempotency — one comment per ticket, ever:

Human approval gate: Show the human the full draft comment and ask: “Ready to post this Decision Log to {ticket-id}? (yes / edit / skip)” — do not post without explicit confirmation.

Identity disclosure (required): The comment body MUST begin with your identity line (as defined in CLAUDE.md) so readers don’t mistake the Decision Log for something the human typed themselves. When overwriting an existing Decision Log comment, refresh this line — do not leave the old one in place.

Comment format:

   ## Decision Log

   _Posted by {your identity} on behalf of @<github-or-jira-handle>._

   _Last updated: YYYY-MM-DD — Push SHA: {short-sha}_

   ### Planning
   - <decision: what was chosen and what was the alternative, e.g. "Chose REST over GraphQL — GraphQL deferred to follow-up">

   ### Implementation
   - <key implementation choice approved by the human>

   ### Review
   - <review outcome, e.g. "Deferred MEDIUM-3 to FOO-456">

   ### Open Items
   - <deferred item> — [FOO-456](link) or "no follow-up ticket yet"

   ⚠️ Pushed with --no-verify — pre-push hooks were bypassed.  ← include only if applicable

Only include sections that have content. Omit empty sections entirely.

If the ticket is a Linear issue:

The Decision Log format is the same as the Jira version (same ## Decision Log header, same sections) — just in Markdown instead of Jira wiki markup.