007af498b8
CI / benchmark-publish (pull_request) Has been skipped
CI / build (pull_request) Successful in 17s
CI / helm (pull_request) Successful in 33s
CI / lint (pull_request) Successful in 3m42s
CI / security (pull_request) Successful in 4m8s
CI / quality (pull_request) Successful in 4m9s
CI / typecheck (pull_request) Successful in 4m20s
CI / integration_tests (pull_request) Successful in 7m2s
CI / unit_tests (pull_request) Successful in 7m55s
CI / docker (pull_request) Successful in 1m19s
CI / coverage (pull_request) Successful in 8m46s
CI / e2e_tests (pull_request) Successful in 16m1s
CI / status-check (pull_request) Successful in 1s
CI / build (push) Successful in 17s
CI / helm (push) Successful in 22s
CI / quality (push) Successful in 31s
CI / lint (push) Successful in 3m28s
CI / typecheck (push) Successful in 3m54s
CI / benchmark-regression (push) Has been skipped
CI / security (push) Successful in 4m5s
CI / integration_tests (push) Successful in 6m13s
CI / unit_tests (push) Successful in 6m28s
CI / docker (push) Successful in 1m34s
CI / coverage (push) Successful in 12m5s
CI / e2e_tests (push) Successful in 18m47s
CI / status-check (push) Successful in 1s
CI / benchmark-publish (push) Successful in 28m0s
CI / benchmark-regression (pull_request) Successful in 59m48s
Renamed all 11 task-type confidence threshold fields in AutomationProfile from phase-transition semantics to spec-defined task-type semantics. Updated all 8 built-in profiles, CLI formatting, YAML schema, services, and all Behave/Robot tests referencing the old field names. Post-review fixes: - Fixed 24 stale old field names in M6 fixture files (automation_profiles.json, autonomy_guardrails.json) - Added model_validator(mode='before') to detect legacy field names and raise actionable ValueError with rename mapping - Added semantic bridge comments in PlanLifecycleService mapping task-type thresholds to phase-transition gates - Added threshold_field to structured log messages for observability - Restored categorised CLI automation-profile show output to match spec (Phase Transitions / Decision Automation / Self-Repair / Execution Controls) instead of flat list - Added missing access_network field to spec show output examples (Rich, Plain, JSON, YAML variants) - Aligned ADR-017 profile fields table to all 11 fields with descriptions matching spec Automatable Tasks table - Aligned automation_profiles.md threshold descriptions with spec - Added spec section references in phase_reversion.md, error_recovery.md, and plan_execute.md for field naming context - Extended repository roundtrip test to assert all 11 threshold fields - Fixed benchmark _make_profile() passing safety fields as top-level kwargs instead of via SafetyProfile sub-model (incompatible with extra="forbid") - Aligned CLI JSON/YAML output structure for automation-profile show with the specification grouped format (phase_transitions, decision_automation, self_repair, execution_controls) - Moved safety boolean fields into the Execution Controls section of Rich output per spec examples - Reverted auto profile description to "Fully automatic except apply" per specification (line 16703, line 28406) - Improved bridge comments in test steps with semantic context for threshold-to-gate mappings ISSUES CLOSED: #902
89 lines
2.4 KiB
Python
89 lines
2.4 KiB
Python
"""Helper script for Robot Framework automation profile smoke tests."""
|
|
|
|
from __future__ import annotations
|
|
|
|
import sys
|
|
from pathlib import Path
|
|
from typing import Any
|
|
|
|
# Ensure src is importable when run from workspace root
|
|
sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "src"))
|
|
|
|
from cleveragents.domain.models.core.automation_profile import (
|
|
BUILTIN_PROFILES,
|
|
AutomationProfile,
|
|
get_builtin_profile,
|
|
)
|
|
|
|
_PROFILE_NAMES = [
|
|
"manual",
|
|
"review",
|
|
"supervised",
|
|
"cautious",
|
|
"trusted",
|
|
"auto",
|
|
"ci",
|
|
"full-auto",
|
|
]
|
|
|
|
|
|
def _test_profile(name: str) -> None:
|
|
"""Load a built-in profile and verify it."""
|
|
profile = get_builtin_profile(name)
|
|
assert isinstance(profile, AutomationProfile)
|
|
assert profile.name == name
|
|
assert 0.0 <= profile.decompose_task <= 1.0
|
|
assert 0.0 <= profile.create_tool <= 1.0
|
|
assert 0.0 <= profile.select_tool <= 1.0
|
|
print(
|
|
f"profile-{name}-ok "
|
|
f"(sandbox={profile.safety.require_sandbox}, "
|
|
f"checkpoints={profile.safety.require_checkpoints}, "
|
|
f"unsafe={profile.safety.allow_unsafe_tools})"
|
|
)
|
|
|
|
|
|
def _test_summary() -> None:
|
|
"""Print summary of all built-in profiles."""
|
|
assert len(BUILTIN_PROFILES) == 8
|
|
for name in _PROFILE_NAMES:
|
|
profile = get_builtin_profile(name)
|
|
thresholds = [
|
|
profile.decompose_task,
|
|
profile.create_tool,
|
|
profile.select_tool,
|
|
profile.edit_code,
|
|
profile.execute_command,
|
|
profile.create_file,
|
|
profile.delete_content,
|
|
profile.access_network,
|
|
profile.install_dependency,
|
|
profile.modify_config,
|
|
profile.approve_plan,
|
|
]
|
|
avg = sum(thresholds) / len(thresholds)
|
|
print(
|
|
f" {name:>12}: avg_threshold={avg:.2f} "
|
|
f"sandbox={profile.safety.require_sandbox} "
|
|
f"checkpoints={profile.safety.require_checkpoints} "
|
|
f"unsafe={profile.safety.allow_unsafe_tools}"
|
|
)
|
|
print("summary-ok")
|
|
|
|
|
|
if __name__ == "__main__":
|
|
cmd = sys.argv[1] if len(sys.argv) > 1 else "summary"
|
|
dispatch: dict[str, Any] = {
|
|
"summary": _test_summary,
|
|
}
|
|
# Add individual profile tests
|
|
for pname in _PROFILE_NAMES:
|
|
dispatch[pname] = lambda n=pname: _test_profile(n)
|
|
|
|
fn = dispatch.get(cmd)
|
|
if fn:
|
|
fn()
|
|
else:
|
|
print(f"Unknown command: {cmd}", file=sys.stderr)
|
|
sys.exit(1)
|