The skills -> command toggle cleanup in
register_enabled_extensions_for_agent() recomputed the remaining
tracked registered_skills names by checking only the toggling agent's
own skills directory. Since registered_skills is a single flat list
shared across every agent an extension has ever been activated under
(skills are only ever rendered for the active agent, so there is no
per-agent registry key), a name whose mirror still existed under a
*different*, previously-active agent's directory was incorrectly
dropped from tracking as soon as the current agent's own copy was
removed. A later full removal only iterates registered_skills, so the
orphaned mirror under the other agent's directory was never found or
cleaned up.
Add _extension_owned_skill_names(), which re-verifies ownership across
every configured agent's skills directory (deduped by shared path) the
same way the existing _unregister_extension_skills() fallback scan
already does, keeping a name only when a SKILL.md with a matching
metadata.source == "extension:<id>" marker is found somewhere -
read-only, no directory creation, no symlink escape. Use it instead of
re-checking only the toggling agent's own directory when recomputing
what remains tracked after narrow stale-mirror cleanup.
Add a red-first regression test: Auggie is activated in skills mode
first (writing a mirror), then Copilot is activated in skills mode
(writing its own mirror for the same names), then Copilot toggles to
command mode. Before the fix, registered_skills lost both names
entirely even though Auggie's mirrors were untouched on disk; after
the fix tracking is preserved and a subsequent full removal correctly
cleans up Auggie's remaining mirrors too.
Assisted-by: GitHub Copilot (model: Claude Sonnet 5, autonomous)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
_infer_legacy_skill_provenance() only probed agents whose registrar
config statically declares extension == "/SKILL.md", excluding
command-backed agents (e.g. Copilot) that can also render preset
overrides as SKILL.md files when ai_skills is enabled. A real
preset-owned .github/skills/.../SKILL.md written while Copilot was the
active skills-mode agent was therefore never probed and got
misattributed entirely to whichever agent activated first after the
upgrade, permanently orphaning Copilot's override on later removal.
Broaden the candidate set to every configured integration
(CommandRegistrar.AGENT_CONFIGS), reusing the existing safe-path
helper (_safe_skills_dir_for_agent, itself built on the shared
_get_skills_dir resolver) rather than inventing new path-construction
logic. The existing preset-marker match (metadata.source ==
"preset:<pack_id>") continues to gate every attribution, so
command-mode agents that never rendered this preset's skill are not
falsely attributed.
Add red-first regression tests: a legacy flat-list entry owned by
Copilot in skills mode, switched directly to Claude with no
intervening Copilot rescaffold, now migrates to a per-agent dict
covering both agents, and removal restores both agents' files instead
of orphaning Copilot's override; plus a negative-case test confirming
a command-mode Copilot with no preset-owned skill marker is not
falsely attributed during the same migration.
Assisted-by: GitHub Copilot (model: Claude Sonnet 5, autonomous)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Three findings from the Copilot review on HEAD b9d9053/3a1e749:
1. `register_enabled_presets_for_agent()` only recorded a preset's command
names into `affected_cmd_names` (the set later passed to
`_reconcile_composed_commands`/`_reconcile_skills`) in the loop that ran
*after* `_register_skills()`, inside the same per-preset `try` block. If
`_register_skills` raised, the `except` caught it and `continue`d before
that loop ever ran — so a preset whose commands phase already wrote real
content to disk never got reconciled against the full priority stack,
leaving its raw content in place instead of a project override or
higher-precedence preset's content. Fix: record the manifest's command
names immediately after the commands phase succeeds and persists, before
calling the independently fallible `_register_skills()`.
2. The legacy flat-list `registered_skills` migration (added for the
previous review round) attributed every name in the list to whichever
agent was currently being (re)activated. If the first operation after
upgrading from a pre-#2948 registry was a direct switch to a *different*
skill-mode agent (e.g. a legacy Claude override, then `integration use
codex` with no intervening Claude rescaffold), the migrated dict only
recorded `{"codex": [...]}`, permanently losing Claude's actual
provenance and orphaning its override on later removal. Fix: added
`_infer_legacy_skill_provenance()`, which probes every configured
skill-mode agent's directory (via the same safe, symlink-validated
helpers already used for restore/removal) for a `SKILL.md` whose
frontmatter records this exact preset as the owner
(`metadata.source == "preset:<pack_id>"`). A name found under more than
one directory is attributed to every matching agent (the preset may have
been active while the user switched between several skill-mode agents
before provenance tracking existed); names that can't be matched to any
directory still fall back to the previously-active best-effort
behaviour. Directory grouping for shared-path aliases (e.g.
agy/codex/zed all resolving to `.agents/skills`) intentionally does not
call `.resolve()` on the path, since doing so diverges from
`project_root`'s own resolution state on platforms where a path
component is itself a symlink (e.g. macOS's `/var` -> `/private/var`)
and made every subsequent containment check spuriously fail.
3. `register_enabled_extensions_for_agent()` has the same command/skill
mutual-exclusion gap the preset path had (fixed in a previous round):
toggling `ai_skills` for the *same active* agent left the opposite
mode's artifact behind. Command -> skills left the extension's
`.agent.md` file and its `registered_commands[agent]` entry in place
once `skills_mode_active` made the commands phase a no-op. Skills ->
command left the extension's `SKILL.md` file in place, since an empty
`_register_extension_skills()` result (because this agent's skills
directory no longer resolves once `ai_skills` is off) was treated as
"nothing to register" rather than "this was rendered here before and is
now stale". This diverges from the preset path in one respect:
`registered_skills` for extensions has always been a flat list with no
per-agent provenance (extension skills are only ever rendered for the
active agent, never per-preset-per-agent tracked), so the fix resolves
ownership by checking which of the extension's tracked skill names
still exist as directories under this specific agent's directory before
removing them — mirroring the same technique `unregister_agent_artifacts`
already uses for full agent deactivation, but scoped narrowly to firing
only when a toggle is actually detected (`skills_mode_active` /
`command_mode_active`), so a same-mode re-run never disturbs
already-correct artifacts or a user's manual customizations.
Regression tests (all confirmed red before their respective fix, green
after):
- tests/test_presets.py::TestPresetSkills::test_rescaffold_reconciles_override_even_when_skills_phase_fails
- tests/test_presets.py::TestPresetSkills::test_rescaffold_legacy_flat_list_direct_switch_preserves_original_agent
- tests/test_extension_skills.py::TestExtensionSkillRegistration::test_rescaffold_toggle_command_to_skills_removes_stale_extension_command_file
- tests/test_extension_skills.py::TestExtensionSkillRegistration::test_rescaffold_toggle_skills_to_command_removes_stale_extension_skill_file
Verification: tests/test_presets.py + tests/test_extensions.py +
tests/test_extension_skills.py (753 passed), tests/integrations/ (1768
passed, 1 skipped), full suite `pytest tests -q` (3942 passed, 109
skipped), `ruff check` on changed files clean.
Assisted-by: GitHub Copilot (model: Claude Sonnet 5, autonomous)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Fix a valid finding from GitHub Copilot's review of HEAD b9d9053 (#2948):
register_enabled_presets_for_agent() normalizes a legacy flat-list
registered_skills value (predating per-agent provenance) to the
{agent_name: [...]} dict shape in memory via _normalize_registered_skills,
but the persistence check only compared the two *normalized* forms. When
the freshly rescaffolded skill names are identical to what the legacy
list already held — the common case, since nothing about the preset or
skill actually changed — that comparison is a no-op and registry.update()
is skipped, leaving the *raw* on-disk value as the un-migrated flat list.
A later switch to a different skill-mode agent and removal then follows
_unregister_skills's legacy best-effort path (restore only the currently
active agent's directory) instead of the per-agent provenance path,
permanently orphaning the first agent's override.
Fix: track the raw (pre-normalization) existing value and force
persistence whenever it's a non-empty list, independent of whether the
normalized content changed. Traced registered_commands for the same
class of bug: its registry value has always been Dict[str, List[str]]
(no legacy flat-list format ever existed for it — the existing
`if not isinstance(existing_commands, dict): existing_commands = {}`
guard is not a lossy migration path), so this fix stays scoped to
registered_skills only.
Added red-first regression test
test_rescaffold_migrates_legacy_flat_list_registered_skills: installs a
preset, overwrites its registry entry with a raw legacy flat list,
rescaffolds the *same* active agent with unchanged skill names, and
asserts the raw registry is migrated to per-agent dict form. Extends the
scenario with a switch to a second skill-mode agent and preset removal
to prove both agents' directories restore cleanly instead of orphaning
the first. Failed against the prior code (raw value stayed a list) and
passes after the fix.
Targeted (tests/test_presets.py, tests/test_extensions.py: 692 passed)
and full (3938 passed, 109 skipped) suites and ruff check on changed
files pass clean.
Assisted-by: GitHub Copilot (model: Claude Sonnet 5, autonomous)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Fix an Important gap in register_enabled_presets_for_agent() surfaced by
quality review (#2948): toggling ai_skills for the *same already-active*
command-backed agent (e.g. `integration upgrade copilot` after flipping
ai_skills, with copilot staying active throughout) left a stale artifact
from the previous mode behind, violating the command/skill mutual-
exclusion invariant this PR otherwise enforces.
- command -> skills: _register_commands()'s ai_skills guard makes the
commands phase a no-op, but the previously-written command file (e.g.
.agent.md) and its registered_commands[agent] entry were never cleaned
up, so it lingered alongside the newly written SKILL.md.
- skills -> command: _get_skills_dir() stops resolving a skills directory
once ai_skills is off, making the skills phase a no-op, but the
previously-written SKILL.md and its registered_skills[agent] entry were
never cleaned up, so it lingered alongside the newly (re)written command
file.
register_enabled_presets_for_agent() now resolves once per call whether
agent_name is a command-backed integration (extension != "/SKILL.md") and
the current ai_skills state, then narrowly unregisters the stale opposite-
mode entry for that agent via the existing _unregister_commands /
_unregister_skills helpers before persisting updated tracking — mirroring
the same per-agent, per-preset isolation already used elsewhere in this
method. Native skill-only agents (claude, codex, ...) are unaffected:
they have no command/skill toggle, so registered_commands and
registered_skills legitimately co-exist for them by design. The trailing
reconciliation pass, project-override precedence, and per-preset
partial-failure isolation are all unchanged.
Added red-first regression tests exercising the real install +
register_enabled_presets_for_agent rescaffold path in both toggle
directions:
- test_rescaffold_toggle_command_to_skills_removes_stale_command_file
- test_rescaffold_toggle_skills_to_command_removes_stale_skill_file
Both failed against the prior code (stale artifact persisted / registry
still tracked it) and pass after the fix.
Targeted (tests/test_presets.py, tests/test_extensions.py: 691 passed)
and full (3937 passed, 109 skipped) suites and ruff check on changed
files pass clean.
Assisted-by: GitHub Copilot (model: Claude Sonnet 5, autonomous)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Fix remaining round-6 review findings on the active-only integration
registration work (#2948):
- register_enabled_presets_for_agent(): registered_commands and
registered_skills were merged and persisted together in a single
registry.update() call after both the commands and skills phases ran.
If _register_skills() raised, the per-preset try/except swallowed it
before that update() call was reached, even though _register_commands()
had already written a real command file to disk. That file became
untracked, so preset removal could no longer clean it up.
install_from_directory() already persists registered_commands
immediately after the commands phase, before starting the independently
fallible skills phase; rescaffold now does the same.
- test_presets.py: renamed a misleading claude_dir variable (pointing at
Gemini's command directory) in
test_composed_none_unregister_respects_active_agent to reuse the
existing gemini_commands_dir variable already defined earlier in the
same test.
Added regression test
test_rescaffold_persists_commands_before_fallible_skills_phase:
simulates a skills-phase failure during rescaffold and asserts the
command file already written to disk is still tracked in
registered_commands.
Verified all other round-6 findings (preset active-integration scoping,
preset reconciliation/remove paths, skills-mode switching, override
precedence during rescaffold, skill-subdirectory symlink safety) are
already addressed by prior commits in this branch; re-checked each
against current code before concluding no further change was needed.
Targeted (tests/test_presets.py, tests/test_extensions.py: 689 passed)
and full (3935 passed, 109 skipped) suites and ruff check on changed
files pass clean.
Assisted-by: GitHub Copilot (model: Claude Sonnet 5, autonomous)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Fix 4 issues from round-6 review of the active-only integration
registration work (#2948):
- remove(): removed_cmd_names only collected primary command names from
registered_commands + manifest aliases, missing commands that were
only ever registered via skills mode (ai_skills guard returns no
command names for command-backed integrations in skills mode). This
skipped reconciliation entirely when removing a higher-priority
skills-mode preset, causing _unregister_skills() to fall back to
core/extension content instead of the surviving lower-priority
preset's override. Now every command template's primary name is
added to removed_cmd_names unconditionally.
- _reconcile_composed_commands(): the "composed is None" branch (fires
when no replace-strategy layer remains for a command, e.g. after
removing a wrap/append preset's base) called unregister_commands()
across every configured non-skill agent, ignoring only_agent. This
deleted historical artifacts from integrations that were never active
for the preset. Now filtered by only_agent like the rest of the file.
- Added _validate_skill_subdir() helper (reusing
_ensure_safe_shared_directory/_validate_safe_shared_directory from
shared_infra.py) and applied it at every site that reads or writes an
individual skill subdirectory (_register_skills,
_unregister_skills_in_dir, _reconcile_skills' override_skills
restoration loop). _safe_skills_dir_for_agent only validated the
parent skills directory; a symlinked leaf subdirectory (e.g.
.claude/skills/speckit-specify) would slip past that check since
is_dir()/exists() follow symlinks, letting write_text/rmtree operate
through it to an arbitrary location outside the project.
Added regression tests: removing a higher-priority skills-only preset
restores the surviving lower-priority preset's content; composed-is-None
unregistration only touches the active agent; symlinked skill subdirectory
rejected on restore; symlinked skill subdirectory rejected on write.
Targeted (934) and full (3934 passed, 109 skipped) test suites and ruff
check pass clean.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- _init_options.py: resolve_active_agent_for_registration() now treats a
dangling init-options.json symlink as present (path.is_symlink() check
alongside path.exists()), since Path.exists() follows symlinks and
returns False for a broken one. Previously a broken symlink fell back
to the legacy "no file" path and registered every detected agent
instead of failing closed.
- presets/__init__.py (register_enabled_presets_for_agent): the
integration use/switch rescaffold path now collects affected command
names across all presets processed and runs
_reconcile_composed_commands/_reconcile_skills once after the loop,
matching install/remove. Previously rescaffolding wrote each preset's
raw content directly with no follow-up reconciliation, so a
project-level override (the highest-priority layer) could be clobbered
by a lower-precedence preset after switching agents.
- presets/__init__.py (_unregister_skills): multiple integrations can
share one physical skills directory (agy/codex/zed all resolve to
.agents/skills). Provenance restoration now groups recorded agent
entries by resolved directory and restores each physical directory
exactly once, preferring the currently active agent's renderer when it
owns that directory (otherwise any recorded owner, chosen
deterministically). Previously each recorded agent key triggered its
own restore pass against the same directory, with whichever agent was
iterated last silently winning regardless of which agent was active.
Adds regression tests for each: a dangling init-options.json symlink
failing closed for both preset resolution and extension add; integration
use rescaffold preserving a project override over a lower-priority
preset; and a codex/agy shared-directory removal restoring the directory
exactly once in the active agent's format.
Targeted (tests/integrations/test_integration_subcommand.py,
tests/test_presets.py, tests/test_extensions.py,
tests/test_extension_skills.py,
tests/integrations/test_integration_opencode.py,
tests/integrations/test_integration_claude.py): 930 passed.
Full suite: 3930 passed, 109 skipped.
ruff check: clean on files touched by this change.
Refs #2948
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Replace the "enumerate every skill-mode directory and restore all of them"
approach from the previous round with precise per-agent provenance
tracking, per reviewer feedback that the enumerate-and-restore-everything
design was unsound:
- registered_skills changes from a flat List[str] to Dict[str, List[str]]
(agent name -> skill names actually written), mirroring the shape
registered_commands already uses. _register_skills now returns this
per-agent mapping instead of a bare list, and every call site
(register_enabled_presets_for_agent, install_from_directory, the
_reconcile_skills "was this skill previously managed" check) is updated
to read/merge the new shape. Legacy flat-list registry entries from
before this change are still readable: writes self-migrate the format,
and _normalize_registered_skills() handles the transitional read paths.
- _unregister_skills now restores exactly the agent directories recorded
for a preset instead of guessing at every skill-mode integration that
happens to exist on disk. This fixes two problems with the old
enumerate-everything design: (1) it could silently overwrite or delete
another preset's (or a user's) override in an agent directory the
current preset never actually touched, and (2) it depended on
transient per-process integration state (_skills_mode), which is unset
in a fresh CLI invocation for mode-selectable integrations like Copilot
--skills, permanently orphaning their overrides after a process
restart. Registries written before this change (flat list, no agent
provenance) fall back to best-effort restoration under only the
currently active agent, matching the pre-existing guarantee level.
- Every directory resolved from persisted provenance is now validated
through the project's shared symlink/containment guard
(_ensure_safe_shared_directory) before any file in it is read, written,
or removed, since restoration may target an agent that isn't currently
active and its directory can't be assumed safe just because a name was
recorded for it.
- _tracked_skill_agent_dirs() (the enumeration helper introduced last
round) is removed; it's superseded by the provenance-based design.
Adds regression tests: a symlinked skills directory is rejected during
removal; removing one preset does not disturb a different preset's
override in another agent's directory; and a Copilot --skills
registration installed, then removed after switching agents in a fresh
PresetManager instance (simulating a new process), is still correctly
restored. Updates existing skill-registration assertions across
test_presets.py and test_integration_claude.py for the new per-agent
registry shape.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Fixes five deeper active-only registration bugs surfaced by Copilot review
after 2486c08, all in the presets/extensions single-active integration
rule (#2948):
1. presets: _reconcile_composed_commands (run after install/remove)
bypassed the active-only filter entirely, writing composition-winner
command files for every detected non-skill agent via
register_commands_for_non_skill_agents. Added an only_agent param to
that registrar method (mirroring register_commands_for_all_agents)
and threaded it through all 5 reconciliation call sites.
2. presets: `integration use copilot` with --skills (ai_skills: true)
wrote both the static .agent.md command file AND the SKILL.md
mirror for the same override. Mirrored the extension path's
ai_skills guard in both _register_commands and the reconciliation
pass: a command-backed active agent running in skills mode is
excluded from non-skill command registration.
3. presets: registered_skills was a flat list, so switching between
two skill-mode agents (e.g. Claude -> Codex) and then removing the
preset only restored the currently active agent's directory,
permanently orphaning the other. _unregister_skills now restores
every existing skill-mode agent directory instead of only the
active one.
4. extensions: load_init_options() collapses "no file" and "corrupted
file" into the same {}, so the round-2 fail-closed fix didn't
actually distinguish them. Added a shared
resolve_active_agent_for_registration() helper in _init_options.py
that checks file existence separately from parse success, returning
a distinct sentinel for "file absent" vs None for "corrupted or
invalid". extensions/__init__.py now uses this helper.
5. presets: same corruption-collapsing bug in _register_commands's
active_agent resolution. Now uses the same shared helper as (4).
Adds regression tests for all five: reconciliation active-only
filtering, copilot --skills dual-write prevention, multi-skill-agent
switch+remove, and corrupted init-options fail-closed behavior for
both extension add and preset add. Each test was verified to fail
against the pre-fix code and pass with the fix.
Targeted (883) and full (3923 passed, 109 skipped) suites pass; ruff
check clean.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- register_enabled_presets_for_agent now processes presets in reverse
priority order (lowest-precedence first) so the highest-precedence
preset is written last and actually wins after `integration use`
rescaffolds two overlapping preset command overrides. Verified this
reproduces the previously reported reversed-priority bug and that the
fix resolves it.
- _register_commands_for_active_agent now checks for the "ai" key's
presence separately from its value: a missing key still falls back to
detection-based registration for all agents, but a recorded, malformed
value (non-string or empty, e.g. [] or null) now fails closed
(registers nothing) instead of being treated as "no active
integration" or reaching AGENT_CONFIGS.get() with an unhashable key
and raising TypeError.
- Updated docs/reference/presets.md and docs/reference/integrations.md
to describe active-only preset/extension registration and clarify
that `integration use`/`switch` is the activation point for
installed extensions and presets, and that `upgrade` only
re-registers them for the active integration.
Adds regression tests: two enabled presets overriding the same command
with different priorities (priority winner must survive `use`
rescaffolding), and a malformed recorded `ai` value ([]) for
`extension add`.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- Restrict the extension-add active-integration fallback to projects
with no recorded active key at all. A recorded but unsupported key
(e.g. "generic", deliberately excluded from AGENT_CONFIGS) no longer
falls back to registering every detected agent.
- Apply the same single-active rule to preset command overrides:
PresetManager._register_commands now scopes registration to the
active integration via only_agent.
- Add PresetManager.register_enabled_presets_for_agent, mirroring
ExtensionManager.register_enabled_extensions_for_agent, and call it
from integration use/switch/upgrade (active only) alongside the
existing extension re-registration so presets are rescaffolded on
activation instead of being written for inactive integrations.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
extension add registered commands for every detected agent, and
integration upgrade back-filled enabled extensions for non-active
integrations. Maintainer direction on #2948: treat the project as
single-active. Only the active integration gets extension artifacts;
use/switch rescaffold the target when the user selects it.
- extension add now routes through the all-agents pass restricted to
the active integration (only_agent), keeping detection and
missing-skills-dir recovery safeguards. Projects without recorded
init-options fall back to detection-based registration.
- integration upgrade re-registers extensions only when upgrading the
active integration, reversing the #2886 back-fill for non-active
targets at maintainer request.
Fixes#2948
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
adapters._validate_remote_url reads parsed.hostname, which raises ValueError
on a malformed authority (e.g. an unclosed ipv6 bracket https://[::1). the
function's contract is to raise BundlerError for any bad url - every other
reject path does - so the raw ValueError leaked to the caller and crashed the
fetch instead of failing cleanly. wrap the parse and convert to BundlerError.
bundler sibling of #3369, which fixed the cli extension/preset/workflow add
paths but not this validator. added a regression test that fails pre-fix.
* fix(workflows): validate scalar types before string operations in workflow validation
YAML parses unquoted scalars like version: 1.0 and id: 123 as
float/int, which crashed validate_workflow and workflow add with raw
tracebacks. Type-check id, name, version and step ids before regex
and string operations so these surface as validation errors. Accept
an unquoted schema_version: 1.0 instead of printing a self-identical
rejection message.
Fixes#3420
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(workflows): treat falsey non-strings as type errors, not missing fields
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(workflows): only accept schema_version 1.0 so the error message is accurate
The check also accepted "1" while the error said Expected '1.0'.
Unquoted YAML 1.0 still works via str(); plain 1 is now rejected with
the message that matches.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The note told readers to "See .specify/templates/plan-template.md for
the execution workflow" — that path is the file itself. The execution
workflow lives in the plan command's definition, which the note already
names via the __SPECKIT_COMMAND_PLAN__ placeholder. Point there instead
of at a self-reference.
Fixes#1148
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* chore: bump version to 0.12.10
* chore: begin 0.12.11.dev0 development
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
The plan command's outline listed two bullets labeled Phase 1 and the
completion report said the command ends after "Phase 2 planning",
but the Phases section only defines Phase 0 and Phase 1. Phase 2
(tasks.md) belongs to the tasks command, as plan-template.md states.
The duplicated "Phase 1: Update agent context" bullet is a leftover
from before the agent-context extension: core no longer ships an agent
script, and the update runs via the extension's after_plan hook.
Fixes#1036
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
`-Number` defaults to 0, so the previous `-eq 0` / `-ne 0` checks could
not distinguish an unset flag from an explicit `-Number 0`: a user
requesting branch `000-...` was silently routed into auto-detection.
Switch both checks to `$PSBoundParameters.ContainsKey('Number')`, which
tests whether the flag was actually supplied — mirroring the bash twin's
empty-string sentinel (`[ -z "$BRANCH_NUMBER" ]` / `[ -n ... ]`).
Add a parity regression test to both TestCreateFeatureBash and
TestCreateFeaturePowerShell asserting `--number 0` / `-Number 0` yields
`000-zero`.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
test_template_renders_python_invocation monkeypatches shutil.which to
return /usr/bin/python3, but on Windows resolve_python_interpreter guards
the which() result with a real _interpreter_runs subprocess probe (#3304).
The mocked /usr/bin/python3 path does not exist on a Windows runner, so the
probe fails, the resolver falls back to sys.executable (a ...python.exe
path), and the python3-anchored regex assertion fails.
Patch _interpreter_runs to return True in the _pin_interpreter fixture so
the resolved interpreter token stays python3 across all platforms, keeping
the #3304 production guard intact while making the assertion deterministic.
Assisted-by: GitHub Copilot (model: Claude Opus 4.8, autonomous)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
* feat(workflows): make shell step timeout configurable
The shell step hardcoded a 300s subprocess timeout, killing any
legitimate long-running QA command. Read an optional timeout field
(seconds, positive integer, default 300) and validate it.
Fixes#3327
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* style: multi-line timeout validation, assert status in default test
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Guard against unvalidated timeout in ShellStep.execute()
The engine does not auto-validate step config, so a string or null
timeout would reach subprocess.run() and crash the run with a
TypeError. Fall back to the 300s default for malformed values,
mirroring how the engine treats unvalidated continue_on_error.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: find plans in nested spec directories
The agent-context mtime fallback used a one-level specs/*/plan.md
glob, so scoped layouts (specs/<scope>/<feature>/plan.md via
SPECIFY_FEATURE_DIRECTORY) never matched and the SPECKIT block was
written without a plan path. Recurse in both script variants and
update the command doc wording.
Fixes#3024
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* refactor: pick newest plan with max(), align doc wording
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* feat(templates): add py: lines to command templates' scripts frontmatter
Every templates/commands/*.md with a scripts: block now declares a py:
variant so --script py renders a Python invocation via the existing
interpreter-prefixing in process_template.
Fixes#3283
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: install scripts/python for the py script type
install_shared_infra mapped every non-sh script type to powershell, so
--script py rendered invocations pointing at files that were never
installed. Map py to the python variant dir and skip __pycache__
artifacts during the copy.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* test: read templates with explicit utf-8 encoding
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: keep py: lines standalone, enforce script existence in tests
Drop the plan/tasks py: lines that referenced scripts shipping in the
core port (#3280); they move to that PR. Tests now assert every py:
line points at a script the repo ships, so a dangling reference can
never merge green.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* chore: bump version to 0.12.9
* chore: begin 0.12.10.dev0 development
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(integrations): skip Windows Store python3 alias stub in resolve_python_interpreter
On stock Windows, python3 on PATH is the Microsoft Store App Execution
Alias stub: it exists but only prints an installer hint and exits
non-zero, so generated {SCRIPT} invocations for the py script type were
broken. Verify the found interpreter actually runs before accepting it,
on Windows only, mirroring the parse-success-not-availability approach
of #3312/#3320 for the sh scripts. POSIX keeps the plain existence
check. sys.executable remains the fallback and is always live.
Fixes#3383
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(integrations): probe interpreter isolated and without site, discard I/O
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* test: pin POSIX platform in PATH-resolution tests
The tests fake shutil.which with POSIX paths; on Windows CI the real
sys.platform made the stub probe run against those fake paths and
fall through to sys.executable.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
yaml.safe_dump with default_style='"' replaces the hand-rolled quote
helpers in base.py and hermes, so newlines and control characters in
template descriptions round-trip instead of producing unparseable YAML.
Fixes#3391
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(workflows): apply chained expression filters left-to-right
The pipe-filter parser in `_evaluate_simple_expression` split the
expression only at the *first* top-level `|` and treated the whole
remainder as a single filter. So a filter chain like
`{{ inputs.rows | map('name') | join(', ') }}` handed
`map('name') | join(', ')` to one filter, where the `(\w+)\((.+)\)`
regex mangled it and raised `ValueError`.
This broke the canonical use of `map`: it returns a list, and `join`
is the only filter that renders a list to a string, so the two are
meant to be chained. Chaining was impossible for every registered
filter.
Split the pipe segments at the top level (quote/bracket aware, so a
literal `|` inside a quoted argument like `join(' | ')` is preserved)
and fold each filter over the running value. The single-filter logic
is extracted verbatim into `_apply_filter`, so all existing strict
handling (`from_json` arity, unsupported-form vs unknown-filter
messages) is unchanged and now applies to every link in the chain.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Potential fix for pull request finding
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
* fix(scripts): resolve invoke_separator by parse success, add awk fallback (#3304)
`get_invoke_separator` in common.sh selected its JSON parser by tool
*availability* (`command -v python3`) rather than parse *success*, and had
no text fallback after the python3 branch. On stock Windows + Git Bash —
no jq, and `python3` resolving to the Microsoft Store App Execution Alias
stub that passes `command -v` but exits 49 at runtime — it silently fell
back to "." even for `-`-separator integrations (e.g. forge, cline). The
observable result was wrong command hints in error messages, such as
`/speckit.plan` instead of `/speckit-plan`, in check-prerequisites.sh and
setup-tasks.sh.
This is the same existence-vs-runtime pattern fixed for the feature.json
parser in #3304's primary report; the reporter explicitly asked that other
python3 call sites be checked. The two remaining sites (resolve_template,
resolve_template_content) already fall through on failure and are unaffected.
- Restructure the jq -> python3 chain to fall through on parse failure,
gated on a `parsed` success flag rather than exclusive elif branches.
- Make the python3 branch signal failure (sys.exit(1)) instead of printing
"." so a stub failure falls through instead of being accepted.
- Add an awk text fallback (portable, no gawk-only whole-file slurp) that
reads the active integration key and its invoke_separator, handling both
pretty-printed (the written form) and compact JSON. Malformed input
safely defaults to ".".
Regression test simulates the broken-stub environment (jq + python3 stubs
that exit 49 on PATH) and asserts the `-` separator is still recovered.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Potential fix for pull request finding
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
* Potential fix for pull request finding
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
* fix(scripts): close awk END block so invoke_separator fallback works
The awk text fallback in get_invoke_separator was missing the closing brace for its END block, so awk aborted with a syntax error on every invocation and the separator silently stayed at the default '.'. On environments with neither jq nor a working python3 (stock Windows + Git Bash, the exact case this fallback exists for) a '-'-separator integration like forge produced a wrong command hint. The bug escaped CI because the test that exercises it is gated on working bash and skips on the Windows dev shell.
Also harden the fallback per review: use the portable '[-.]' character class (a leading '-' inside '[]' can be read as an ill-defined range on some awk builds), and correct the test docstring, which said 'with no jq' though the helper installs a present-but-failing jq stub. Addresses Copilot review feedback on PR #3320.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
* fix(shared-infra): refresh_shared_templates preserves recovered user files
refresh_shared_templates skipped a shared template only when it was
untracked or modified, ignoring the manifest's is_recovered marker that
install_shared_infra already honors. So a pre-existing user template
(adopted via record_existing(recovered=True), hence tracked and
hash-unmodified) was silently overwritten with bundled content on
refresh — the exact data-loss class that #2918 fixed for
install_shared_infra. Add the is_recovered check to the skip predicate.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(shared-infra): include recovered files in refresh skip warning
The skip predicate now also skips recovered (pre-existing user) files,
so the warning saying only 'modified or untracked' could mislead a user
into thinking they edited a file that was simply recorded as recovered.
Reword to 'modified, untracked, or preserved (recovered)'.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix(agents): resolve skill placeholders in Goose (yaml) command output
CommandRegistrar.register_commands resolves {SCRIPT}/__AGENT__ and the
$ARGUMENTS placeholder in the markdown and toml branches, but the yaml
branch (Goose recipes) called render_yaml_command directly, skipping
both. So extension/preset command bodies installed for Goose kept literal
{SCRIPT}, __AGENT__, and repo-relative script paths in the generated
.goose/recipes/*.yaml prompt. Mirror the markdown/toml branches: run
resolve_skill_placeholders + _convert_argument_placeholder on the body
before render_yaml_command.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(goose): assert positive placeholder replacements in recipe prompt
Per review: parse the generated recipe with yaml.safe_load and assert the
prompt contains the resolved values (.specify/scripts/, 'agent goose',
{{args}}), not merely that the literal tokens are absent — a wrong output
that happens to omit the exact strings would otherwise pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix(bundler): enforce version pin on bundled preset/extension installs
The bundled install branch of _PresetKindManager/_ExtensionKindManager
called install_from_directory and returned before _assert_pinned_version,
so a bundle manifest pinning e.g. 2.0.0 would silently install the
bundled asset's own version (1.0.0) — the pin was only enforced on the
catalog path. _WorkflowKindManager already enforces it unconditionally.
Read the bundled asset's declared version from its manifest (best-effort;
None => cannot enforce, matching the catalog 'advertises no version'
escape hatch) and call _assert_pinned_version before install, in both
bundled branches.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(bundler): address review on bundled version-pin check
- Make _assert_pinned_version's error source-agnostic ('resolved version'
/ 'the source') so bundled preset.yml/extension.yml mismatches read
correctly, not just catalog ones.
- Type-guard _bundled_manifest_version: only a non-empty string version is
usable; missing/non-string/whitespace -> None ('cannot enforce').
- Add bundled-preset success-path test (matching pin + version=None both
proceed to install_from_directory), mirroring the extension test.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* test: isolate integration test home
Assisted-by: Codex (model: GPT-5, autonomous)
* test: assert integration home isolation
Assisted-by: Codex (model: GPT-5, autonomous)
* test: extend integration home isolation to module-scoped setup
Address Copilot review on #3144.
Add a session-scoped autouse fixture so HOME/USERPROFILE/XDG are redirected
for setup that runs outside a test function (e.g. the module-scoped status_*
fixtures in test_integration_subcommand.py that run `specify init` before any
per-test isolation applies). The function-scoped fixture still overrides HOME
per test.
Also assert Path.home() resolves to the isolated home, since most integrations
(Hermes, catalog) read the home via that API rather than the env vars directly.
* chore: bump version to 0.12.8
* chore: begin 0.12.9.dev0 development
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* Add LLM Wiki extension to community catalog
Add wiki extension submitted by @formin to:
- extensions/catalog.community.json (alphabetical order)
- docs/community/extensions.md community extensions table
Closes#3319
Assisted-by: GitHub Copilot (model: claude-sonnet-4.6, autonomous)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: limit catalog.community.json changes to wiki entry + timestamps only
Reverts the unintended reordering and reformatting of existing extensions
(aide, checkpoint, critique, threatmodel, etc.) and companion's tools array.
Only the new wiki entry and updated_at timestamps are now changed.
Assisted-by: GitHub Copilot (model: claude-sonnet-4.6, autonomous)
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
* docs: document missing flags and integrations
* docs: remove invalid --refresh-shared-infra from upgrade command
* docs: address PR feedback for extension and integration flags
* docs: reorder extension add options to match CLI help
* feat(extensions): port update-agent-context to Python
Ports the agent-context extension updater to a single Python script,
per #3281 and the check-prerequisites PoC pattern from #3302. The bash
version already ran its core logic through embedded Python heredocs, so
the port lifts that logic into a standalone script. Parity tests run
bash and Python side by side and compare output and resulting
context-file bytes.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(extensions): match bash case-insensitivity on MSYS, test unparseable config gate
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(scripts): fall through to grep/sed when python3 is a broken stub in feature.json parser
read_feature_json_feature_directory picked its json parser by availability
(if jq / elif python3 / else grep-sed). on windows `python3` usually resolves
to the microsoft store app execution alias stub: it satisfies `command -v` but
fails at runtime (exit 49). the elif selected it, the runtime failure was
swallowed to _fd='', and the grep/sed last resort was never reached - so a
valid .specify/feature.json read as empty and every setup-plan / setup-tasks /
check-prerequisites call errored with "Feature directory not found" right after
a successful `specify init --script sh`.
change selection from availability to parse success: try jq, then python3 only
if still empty, then grep/sed only if still empty. a parser that exists but
produces nothing now falls through instead of terminating the chain.
the write path (_persist_feature_json) already uses jq-or-printf with no
python3, so it was unaffected; only the read path needed this.
add a regression test that puts a python3 stub (exit 49, like the store alias)
first on PATH and asserts setup-plan.sh still resolves the feature via the
grep/sed fallback.
fixes#3304
* test: shadow jq so the broken-python3 fallback test actually exercises it
the test claimed it dropped jq so the parser chain would reach python3 and
then grep/sed, but it only prepended the python3 stub dir to PATH. on a
runner with jq installed, read_feature_json_feature_directory parses via jq
and never reaches the fallback the test is meant to cover.
add a failing jq stub alongside the python3 stub so the chain is forced
through jq -> python3 -> grep/sed regardless of what the runner has installed.
* fix(toml): escape control characters so generated command files parse
both toml renderers (TomlIntegration._render_toml_string for gemini/tabnine
and CommandRegistrar.render_toml_command for extension/preset commands) wrote
control characters raw into a multiline or basic string. toml forbids literal
control chars (U+0000-U+001F except tab/newline, and U+007F) in every string
form, and a bare CR that is not part of a CRLF pair, so a description or body
containing one produced a .toml file that fails to parse.
route any value with such a character to a fully-escaped basic string that
emits the leftover control chars as \uXXXX. added regression tests that
round-trip through tomllib.
* refactor(toml): centralize control-char escaping in one shared helper
the control-char detection and basic-string escaping added for both toml
renderers were copy-pasted into agents.py and integrations/base.py. move the
two functions into specify_cli/_toml_string.py and have both renderers
delegate to it, so the escaping rules can't drift apart later.
no behavior change; both renderers now reference the same implementation.
* fix(cli): exit cleanly on malformed IPv6 URLs in extension/preset/workflow add
extension add --from, preset add --from, and workflow add <url> parsed
the user-supplied URL with a bare urlparse before their HTTPS/host
validation, so an unclosed IPv6 bracket escaped as a raw ValueError
traceback. Wrap each parse and emit the surrounding validation's clean
error style + typer.Exit(1) instead.
Fixes#3368
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(cli): convert malformed redirect URLs to URLError in shared redirect handler
Parse the redirect target once in _StripAuthOnRedirect.redirect_request
before the validator and stdlib handler run, converting ValueError into
URLError which every download path already catches. Also escape from_url
in the preset install message so IPv6 brackets don't break Rich markup.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
resolve_github_release_asset_api_url's is_ghes branch built the authority
with 'parsed.port', which raises ValueError on a malformed port (e.g.
host:notaport). The function's contract is to resolve or return None,
never raise — every other unresolvable case returns None. An allowlisted
GHES host with a bad port therefore crashed the caller. Read parsed.port
defensively and return None on ValueError.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
manifest.py::_sha256 does an unguarded open(). check_modified() and
uninstall() both call it on a readable-but-unopenable regular file
(e.g. permission denied) without catching OSError, so
'specify integration upgrade/uninstall/switch' surface a raw
PermissionError traceback. Guard both call sites: in check_modified()
treat an unreadable file as modified (consistent with the adjacent
symlink / non-regular-file handling); in uninstall() treat it as skipped
and preserve it (mirroring the existing path.unlink() OSError guard just
below). The force short-circuit is unchanged.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* chore: bump version to 0.12.7
* chore: begin 0.12.8.dev0 development
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>