* skillopt_sleep/evidence.py — append-only, thread-safe, redacted
evidence.jsonl per night: harvest sessions -> miner exchanges (verbatim
prompt/reply) -> mined tasks with checks -> split assignment -> every
replay attempt (phase-tagged, cache hits marked) -> per-task scores with
failing checks named -> reflect exchanges + parsed edits -> gate trials
and the final decision with its score arithmetic -> staged artifacts.
Config-gated (evidence_log, default on; evidence_max_chars cap).
* skillopt_sleep/prompts.py — central registry of the four LLM prompt
templates (miner/attempt/judge/reflect), byte-identical defaults to the
previously inlined strings; user overrides in prompts.json take effect
on the next model call (mtime-checked), no restart needed.
* cycle.py builds dual backends from config (optimizer_*/target_*),
pre-creates the staging dir so evidence lands beside the report;
staging.new_staging_dir() de-collides same-second runs;
latest_staging() now skips non-adoptable (evidence-only) folders.
* tests/test_sleep_evidence.py — 8 no-network stdlib tests: chain
completeness, redaction/truncation/ordering, disable flag, prompt
override round-trip + live effect, no-tasks-night adoption guard.
The Devin MCP server's backend enum only listed mock/claude/codex/copilot,
excluding the handoff backend that was merged to the engine in #125. This
made the subscription-friendly, no-API-key path unavailable to Devin users
while Claude Code had a dedicated /skillopt-sleep-handoff command for it.
Add "handoff" to the backend enum in _TOOL_SCHEMA so sleep_run accepts
backend: "handoff". The engine already handles the prompt/answer file loop
(exit code 3 + .skillopt-sleep-handoff/); the MCP server needs no special
handling — it passes through the engine output showing pending prompts.
Update the README, rules snippet, and test_backends_in_enum accordingly.
Addresses maintainer review on the OpenAI-compatible endpoints PR:
1. CLI: accept --backend azure_openai in skillopt_sleep/__main__.py (the
documented command was rejected by the argparse choices).
2. Example runner: exit with the child's return code so watchdog/supervisors
see a failed sleep run as a failure.
3. Error state: clear last_call_error when a retry recovers; set an explicit
"empty response on all N attempts" diagnostic when every attempt returns
empty text.
4. Security guard: the managed-identity path now refuses to send an Azure AD
bearer token to any endpoint outside *.openai.azure.com /
*.cognitiveservices.azure.com — a custom endpoint requires explicit
AZURE_OPENAI_AUTH_MODE=openai_compatible + API key.
5. Provider-neutral requests: compat mode sends only the standard contract
(max_tokens, default 8192 via SKILLOPT_SLEEP_COMPAT_MAX_TOKENS);
provider-specific body fields are opt-in via SKILLOPT_SLEEP_CHAT_EXTRA_BODY
(JSON) — the deepseek model-name inference is removed.
6. Docs: removed the unimplemented OPTIMIZER_*/TARGET_* env-var claim; added a
configuration reference matching the implementation exactly.
7. Tests: tests/test_azure_openai_compat.py — 17 deterministic no-network
unittest cases covering CLI acceptance, compat-vs-Azure client selection,
endpoint resolution, the credential guard, request kwargs (opt-in extra
body / token cap), retry-success error clearing, empty-response
diagnostics, and runner exit-code propagation.
Re-verified live against DeepSeek (deepseek-v4-pro, openai_compatible mode)
after the rework: client type OpenAI, completion returned, no error state.
The LinearScheduler.known_decay_sequence test now reflects the new
t = (step-1)/(total_steps-1) contract where step 1 returns max_lr.
Updated docstrings for midpoint and early-step tests to match.
All 43 tests now pass against the updated formula.
The qwen_chat backend (the generic OpenAI-compatible client) hardcoded
max_tokens and always sent temperature, so reasoning models behind
OpenAI-compatible gateways (GPT-5.x, Claude Opus 4.8 via Azure/LiteLLM)
would 400.
- Add opt-in QWEN_CHAT_USE_MAX_COMPLETION_TOKENS (+ role variants) that
swaps the payload key max_tokens -> max_completion_tokens.
- Treat an explicit empty / none / off temperature as "omit" instead of
collapsing to the 0.7 default (via _resolve_temperature).
- Thread both through configure_qwen_chat / _update_config.
- Defaults unchanged; fully backward compatible. Adds 6 tests.
Fixes#127
Co-authored-by: Chirag Singhal <chirag127@users.noreply.github.com>
Adds --backend handoff: the engine runs all deterministic stages and
outsources attempt/judge/reflect to prompt/answer files an interactive
agent session fills between runs (exit 3 = pending batch, re-run to
resume). Deterministic replay + the prompt-hash answer cache make resume
stateless; sentinel detection aborts any call built from unanswered
output so placeholders never reach scores or staging. Session digests
and mined tasks are pinned per night (secret-redacted) so the sessions
answering prompts cannot shift the task set, and LLM mining is routed
through the same handoff files. Ships a /skillopt-sleep-handoff Claude
Code command that answers each prompt in a fresh-context subagent to
protect the held-out gate.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix(skillopt-sleep): surface Claude CLI spawn failures instead of silent zero scores
In _call and attempt_with_tools, the bare 'except Exception: return ""'
swallows FileNotFoundError (e.g. bare 'claude' on Windows with npm .cmd
shim) and any other spawn failure, returning an empty string that the
trainer treats as a legitimate model response that scores 0.0 everywhere.
Now the exception is caught explicitly: last_call_error is set, a
warning is logged, and the empty-string return is preserved for
backward compatibility of the control flow. This mirrors the pattern
from #92 (codex backend) which fixed the same class of 'dead CLI
masquerades as nothing to learn' bug.
Issue: #121
* test(skillopt-sleep): add tests verifying Claude CLI spawn failures are surfaced
Add two tests to TestClaudeCliBackendBare:
- test_spawn_failure_sets_last_call_error: _call sets last_call_error
and returns '' when subprocess.run raises FileNotFoundError.
- test_attempt_tools_spawn_failure_sets_last_call_error: same for
attempt_with_tools.
These prove the fix from the parent commit (surface spawn failures
instead of silently scoring 0) and guard against regressions.
Add explicit assertions for held-out scores and gate actions to the verifier discipline test suite to strengthen its guarantees.
- Assert the concrete held-out baseline and candidate scores in test_gate_rejects_reward_hacking_edit.
- Add test_gate_accepts_beneficial_edit using MockBeneficialBackend to provide a paired case where an edit genuinely improves the held-out slice, expecting accepted=True and gate_action='accept_new_best'.
The Claude Code and Codex plugin shells' runner (run-sleep.sh) only
resolved the engine by searching for a skillopt_sleep/ source directory
on disk. After v0.2.0 shipped the engine on PyPI (skillopt-sleep CLI),
the plugins still errored for users who installed via pip/uv without
cloning the repo — even though the error message said "pip install
skillopt" would work.
Add two fallbacks before the error:
1. skillopt-sleep CLI on PATH (covers uv tool install, pipx, pip install)
2. python -m skillopt_sleep when importable (covers pip install into
the active Python)
Fallback 1 is checked before fallback 2 because uv tool install / pipx
isolate the package from the system Python's import path, so the import
check would fail even though the CLI is available.
The existing source-checkout resolution path is unchanged — repo-clone
users keep working as before. The fix lands in both plugins/run-sleep.sh
(shared, used by Codex) and plugins/claude-code/scripts/run-sleep.sh
(bundled copy, byte-identical, used by the Claude Code marketplace
install which fetches only the plugins/claude-code/ subtree).
Runnable SearchQA splits should remain disjoint. A duplicate manifest id across train, val, or test previously collapsed into the wanted-id set and reused the same row in multiple output splits without warning.
Constraint: Preserve manifest order and output schema for valid manifests.
Rejected: Deduplicate automatically | hiding split overlap would make evaluation contamination harder to notice.
Confidence: high
Scope-risk: narrow
Reversibility: clean
Directive: Fail fast on split-manifest integrity problems instead of repairing them silently.
Tested: uv run --with pytest pytest tests/test_materialize_searchqa.py -q
Tested: uv run --with ruff ruff check scripts/materialize_searchqa.py tests/test_materialize_searchqa.py
Not-tested: Loading the live Hugging Face dataset.
Rollout hard scores can be continuous when smoothed rewards are used. Converting the field through int() turned values like 0.75 into 0, which loses signal before scoring and serialization.
Constraint: Keep the existing RolloutResult dictionary shape unchanged.
Rejected: Clamp hard scores to 0 or 1 | contradicts existing continuous-score support in compute_score and sleep replay types.
Confidence: high
Scope-risk: narrow
Reversibility: clean
Directive: Treat hard as numeric reward data, not only a binary label.
Tested: uv run --with pytest pytest tests/test_types.py tests/test_scoring.py -q
Tested: uv run --with ruff ruff check skillopt/types.py tests/test_types.py
Not-tested: Full benchmark rollouts.
LLM responses can include bracketed prose before the actual JSON array. The old greedy regex spanned from the first '[' to the last ']', causing valid single-array answers to be dropped and making multiple arrays indistinguishable.
Constraint: Keep object extraction behavior unchanged.
Rejected: Return first parseable array | silently guesses when a response has multiple valid arrays.
Confidence: high
Scope-risk: narrow
Reversibility: clean
Directive: Keep array extraction conservative when more than one valid top-level array is present.
Tested: uv run --with pytest pytest tests/test_json_utils.py -q
Tested: uv run --with ruff ruff check skillopt/utils/json_utils.py tests/test_json_utils.py
Not-tested: Full benchmark rollouts.
PR #92 added a per-cycle diagnostics.json that surfaces backend stderr,
optimizer replies, and task responses so a 0.0 night is self-diagnosing.
Those free-text fields can carry credentials (e.g. a codex 401 stderr dump
containing an auth token), so persisting them verbatim was a new on-disk
leak surface.
- Add a shared redact_secrets() in staging.py and route diagnostics.json's
call_error / reflect_raw_head / holdout_detail through it before writing.
- Redact the codex and Claude auth-error log lines too (a secondary sink
when a file log handler is attached); last_call_error stays raw in memory
so _AUTH_MARKERS matching is unaffected.
- Centralize _SECRET_PATTERNS in staging.py (harvest_codex now reuses them)
and extend coverage to AWS / GitHub / Slack / Google / JWT token shapes.
- Tests: secret-shape coverage, private-key blocks, recursive/scalar
passthrough, no over-redaction of plain prose, fail-fast auth-error log
redaction, and an end-to-end check that diagnostics.json has no secret.
Observability-only; the gate and learning algorithm are unchanged.
Co-Authored-By: Claude <noreply@anthropic.com>
Splits CodexCliBackend._call into _call_once + a retry wrapper so transient empties/timeouts are retried instead of silently scored 0, and fails fast on fatal auth/model/version errors (401, refresh_token_reused, token_expired, ChatGPT-account-unsupported, newer-Codex-required). On non-zero exit the CLI error text is surfaced via last_call_error instead of being returned as a model response. Adds per-cycle diagnostics.json (observability only; gate and learning algorithm unchanged) so a 0.0 night self-explains.