TEST-INFRA: [ci-pipeline-design] Centralize and manage tool versions #10953
Merged
HAL9000
merged 3 commits from 2026-06-15 06:23:59 +00:00
task/ci-centralize-tool-versions into master
Labels
Clear labels
auto/needs-reevaluation
controller-managed
overdue
auto/blocked-by-deps
auto/ci-timeout
auto/claimed-implementer
auto/claimed-merge
auto/claimed-reviewer
auto/driver-down
auto/invariant-violation
auto/last-attempt-tier-0
auto/last-attempt-tier-1
auto/last-attempt-tier-2
auto/last-attempt-tier-min
Automation Tracking
auto/needs-conflict-resolution
auto/needs-implementer
auto/postmortem
auto/ready-to-merge
auto/restart-throttled
auto/revert
auto/sentinel
auto/stale-inactivity
auto/unstable
Blocked
Needs Feedback
Signed-off: Owner
Signed-off: Scrum Master
Signed-off: Tech Lead
Spike
Controller deferred this PR; awaiting Phase 6+ scope-evaluator or operator re-enablement.
Auto-agents controller manages this PR/issue (see tools/controller/deploy/RUNBOOK.md). Remove this label to abandon controller management.
PR blocked by an open issue dependency. Operator must close the dep (or remove the dependency link) before the merge driver can act. Auto-cleared by merge_drive when no open deps remain.
Most recent merge cycle hit CI timeout. Driver excludes this PR while last merge_cycle row is < 30 min old; label persists thereafter as visible history.
Currently being processed by an implementer worker.
Currently being processed by the merge driver.
Currently being processed by a reviewer worker.
Merge driver heartbeat stale; pipeline halted. Closed automatically on next clean tick.
Detected master commit violating the strict merge invariant. Tracked as an issue (not a PR label); kept here for label completeness.
In-cycle escalation: most recent attempt ran at the Tier 0 slot (`tier-0`). Slot's model defined in .opencode/models/tiers.yaml.
In-cycle escalation: most recent attempt ran at the Tier 1 slot (`tier-1`). Slot's model defined in .opencode/models/tiers.yaml.
In-cycle escalation: most recent attempt ran at the Tier 2 slot (`tier-2`). Slot's model defined in .opencode/models/tiers.yaml. Gated behind IMPLEMENTER_ESCALATION_TIER2_ENABLED.
In-cycle escalation: most recent attempt ran at the Tier -1 slot (`tier-min`). Slot's model defined in .opencode/models/tiers.yaml. Suffix is ``-min`` (not ``--1``) so the Forgejo UI reads naturally.
Tracking issues used by the AI Automation system for agents to communicate and report.
Rebase conflict needs LLM conflict-resolver.
Failing CI needs implementer attention.
Documenting a driver incident or rollback.
Reviewer has APPROVED this PR and no later REQUEST_CHANGES is outstanding. The merge driver requires this label to even consider a PR for merging. Set by the reviewer worker on APPROVE; cleared on REQUEST_CHANGES.
Train repeatedly lost master-tempo races. Driver excludes via merge_cycle until cooldown elapses; label persists as visible history.
Revert PR backing out an invariant violation. Fast-tracked through the merge driver.
Sentinel PR duplicated from upstream into a personal fork by tools/duplicate_prs_to_fork.py for pipeline testing. Lives only in the fork; the canonical pipeline never sees it.
No implementer activity for N days. Flagged for human review. Auto-cleared on next push to head branch.
Repeatedly fails on current master (>= 3 ci-fail-on-rebased-sha releases in 12 h). Excluded from driver until human triage.
A ticket in a blocked state and unable to complete until some other task is completed first.
Bounty
$100
A bounty of $100 for any open-source contributor who provides a MR that solves this issue
Bounty
$1000
A bounty of $1000 for any open-source contributor who provides a MR that solves this issue
Bounty
$10000
A bounty of $10000 for any open-source contributor who provides a MR that solves this issue
Bounty
$20
A bounty of $20 for any open-source contributor who provides a MR that solves this issue
Bounty
$2000
A bounty of $2000 for any open-source contributor who provides a MR that solves this issue
Bounty
$250
A bounty of $250 for any open-source contributor who provides a MR that solves this issue
Bounty
$50
A bounty of $50 for any open-source contributor who provides a MR that solves this issue
Bounty
$500
A bounty of $500 for any open-source contributor who provides a MR that solves this issue
Bounty
$5000
A bounty of $5000 for any open-source contributor who provides a MR that solves this issue
Bounty
$750
A bounty of $750 for any open-source contributor who provides a MR that solves this issue
MoSCoW
Could have
Could have feature in order to satisfy the epic/legendary.
MoSCoW
Must have
Must have feature in order to satisfy the epic/legendary.
MoSCoW
Should have
Should have feature in order to satisfy the epic/legendary.
There are questions in the ticket that can not be completed until the project owner provides clarity.
Points
1
1 man-hours worth of work for an expert with no learning curve.
Points
13
13 man-hours worth of work for an expert with no learning curve.
Points
2
2 man-hours worth of work for an expert with no learning curve.
Points
21
21 man-hours worth of work for an expert with no learning curve.
Points
3
3 man-hours worth of work for an expert with no learning curve.
Points
34
34 man-hours worth of work for an expert with no learning curve.
Points
5
5 man-hours worth of work for an expert with no learning curve.
Points
55
55 man-hours worth of work for an expert with no learning curve.
Points
8
8 man-hours worth of work for an expert with no learning curve.
Points
88
88 man-hours worth of work for an expert with no learning curve.
Priority
Backlog
This ticket has backlogged priority and is not to be worked on yet
Priority
CI Blocker
Critical priority issue that blocks CI/CD pipeline and prevents PR merges
Priority
Critical
The priority is critical
Priority
High
The priority is high
Priority
Low
The priority is low
Priority
Medium
The priority is medium
When an epic or legendary is in review it must be signed off by owner, tech lead, and scrum master before being marked as completed.
When an epic or legendary is in review it must be signed off by owner, tech lead, and scrum master before being marked as completed.
When an epic or legendary is in review it must be signed off by owner, tech lead, and scrum master before being marked as completed.
A ticket for learning a tool or technology that is needed to be able to do future planning and design.
State
Completed
The ticket has been fully implemented, completed, and merged with the source code. This label should only be applied once a ticket is closed.
State
Duplicate
A ticket that represents the same content as an existing ticket.
State
In Progress
A ticket that is actively being developed.
State
In Review
A ticket that has had some code completed to implement but is waiting to pass peer review and is not yet merged in.
State
Paused
This ticket's work started but wasn't finished. It's on hold (likely in a feature branch) and will be resumed later, either due to a blocker or a delay.
State
Unverified
All new tickets start in this state. A developer may set it to show the ticket is unverified. This means we haven't agreed to work on it. It will either move to a verified state or be closed as wontdo.
State
Verified
The issue has been verified by a developer as legitimate. It will be worked on and verified tickets are now considered part of the backlog.
State
Wont Do
This ticket has been decided it wont be done. This may mean the bug has been determined to not be real (cant verify) or the feature is one we have decided we dont want to adopt.
Type
Automation
Any edits or discussion about the AI automated coding system.
Type
Bug
Something that doesnt work as intended.
Type
Discussion
Anytime a ticket represents a discussion about a subject and doesnt fall into one of the other categories.
Type
Documentation
An error or improvement needed in the documentation.
Type
Epic
Any first tier epic. That is, an epic which contains only issues as children and will not have sub-epics.
Type
Feature
Some new functionality not present.
Type
Legendary
A type of Epic which will contain other Epics.
Type
Refactor
A code change that restructures existing code without changing its external behavior.
Type
Support
Someone needs help using the project.
Type
Task
A generic task that doesnt fit into the other type categories.
Type
Testing
Work exclusively focusing on fixing or expanding testing.
Projects
Clear projects
No project
Assignees
aditya (Aditya Chhabra)
aleenaumair (Aleena Umair)
brent.edwards (Brent Edwards)
CoreRasurae (Luis Mendes)
drew (Drew Morris)
eugen.thaci (Eugen Thaci)
freemo (Jeffrey Phillips Freeman)
HAL9000 (HAL 9000)
HAL9001 (HAL9001)
hamza.khyari (Hamza Khyari)
hurui200320 (Rui Hu)
justin.morris
khird (Kyle Hird)
org.cleveragents
Clear assignees
No Assignees
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: cleveragents/cleveragents-core#10953
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "task/ci-centralize-tool-versions"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Summary
Centralized tool version management into a single
.tool-versionsfile to serve as the single source of truth for all tool versions used in the CI pipeline.Changes
.tool-versionsfile containing UV_VERSION, PYTHON_VERSION, HELM_VERSION, and KUBECONFORM_VERSION.forgejo/workflows/ci.ymlto include a newload-versionsjob that reads from.tool-versionsand exports versions as outputsload-versionsand use the exported version outputs instead of hardcoded valuesfeatures/ci_workflow_validation.featureto verify versions are sourced from the central filescripts/load-tool-versions.shfor local version loadingBenefits
.tool-versions)Testing
load-versionsjob exists and outputs the correct versionsload-versionsand use the exported outputsCloses #1918
Automated by CleverAgents Bot
Supervisor: Implementation | Agent: implementation-worker
Review of PR #10953: TEST-INFRA: [ci-pipeline-design] Centralize and manage tool versions
Linked Issue: Closes #1918 — Centralize tool version management into a single source of truth
CI Status (BLOCKING)
CI is failing on 3 required gate checks:
lint— FAILING after 46sunit_tests— FAILING after 9m20scoverage— SKIPPED (due to unit_tests failure)Per company policy, all 5 required CI gates (lint, typecheck, security, unit_tests, coverage) must pass before a PR can be approved and merged. This is the primary blocker.
Blocking Issues
Missing type annotations on all 6 new step functions — The project enforces zero-tolerance for untyped parameters (Pyright strict). Every function signature must have typed
context(type: Any), typed positional arguments, and typed return values.CI lint job is failing — This will be resolved once type annotations are added, as ruff flags untyped parameters (per project conventions). The CI lint check is a required-for-merge gate.
CI unit_tests is failing — Likely cascading from the test infrastructure unable to handle the new step definitions without proper signatures, or the Behave runner encountering errors during scenario execution. Needs investigation and fix.
Missing PR labels — The PR has zero labels. Per checklist requirement #12, exactly one
Type/label is mandatory (e.g.,Type/Taskfor infrastructure work). Additionally, aPriority/label should be applied based on triage guidance.Non-Blocking Observations
Branch naming — The branch
task/ci-centralize-tool-versionsdoes not match any prescribed prefix in the contributing guidelines (which requirefeature/mN-,bugfix/mN-, ortdd/mN-). Chore/infrastructure changes should usefeature/mN-format.Redundant dependency declarations — In
ci.yml, bothcoverageanddockerincludeload-versionsin theirneeds:arrays (e.g.,[load-versions, lint, typecheck, ...]). Since those jobs already depend on all the same sub-jobs that requireload-versions, this is transitively satisfied and unnecessary.Feature file scenario scope — The new BDD scenarios (
Tool versions file exists, etc.) validate YAML file structure (presence of keys, outputs, dependencies in .yml parsing) rather than exercising runtime behaviors. These read more like implementation smoke tests than living documentation. Consider whether these test actual user-facing capabilities or just assert that the files are well-formed.Incomplete job dependency assertions — The scenario "CI workflow jobs depend on load-versions" checks lint, typecheck, unit_tests, integration_tests, e2e_tests, helm, and build — but does not assert push-validation (which also depends on
load-versionsin the diff).Overall Assessment
The architectural approach of centralizing tool versions into
.tool-versionsis sound and directly addresses issue #1918. Theload-versionsjob pattern is reasonable. However, the missing type annotations create a cascade failure across lint → unit_tests → coverage that blocks all required CI gates from passing.Fix the three items above (type annotations, add PR labels) and re-push for a fresh CI run.
@@ -0,0 +2,4 @@# This file is the single source of truth for tool versions used in the CI pipeline.# Format: TOOL_NAME=versionUV_VERSION=0.8.0Good approach. The format is clean: key=value pairs with a comment header explaining the convention. No issues here.
Automated review completed for PR #10953.
Result: REQUEST_CHANGES
Key blockers:
See full review body for detailed findings and non-blocking suggestions.
Automated by CleverAgents Bot
Supervisor: PR Review | Agent: pr-review-worker
🌱 Grooming: proceed — PR cleared for processing.
(check
no_duplicates, categoryno_duplicates)Anchor PR #10953 focuses specifically on centralizing tool version management into a single
.tool-versionsfile, including modifications to CI workflows and comprehensive BDD tests. Scanned 355 open PRs; no duplicate found. Related CI/test-infra PRs (#10954 Dockerfile security, #1618 reusable workflows, #10845–#10959 performance optimizations) address distinct aspects of the CI pipeline. No other PR targets tool version centralization or.tool-versionsfile creation. Topical differentiation is clear.📋 Estimate: tier 1.
Multi-file CI infra change (5 files, +241/-22): new .tool-versions, ci.yml load-versions job, BDD feature scenarios, step definitions, helper script. Two CI failures require cross-file reasoning: (1) ruff format fix on ci_workflow_validation_steps.py is mechanical but (2) two failing BDD scenarios need understanding of how step definitions inspect the ci.yml YAML — the hardcoded python3.13/UV_VERSION references were replaced with ${{ needs.load-versions.outputs.X }} variables, and the step definitions need to be updated to match this new syntax. Requires holding feature file + step definitions + ci.yml context simultaneously. Standard tier-1 engineering work.
(attempt #4, tier 1)
🔧 Implementer attempt —
rebase-failed.Blockers:
c9c4f154c7to95b816dd44d0a0b450a2tod8b4823b75(attempt #30, tier 2)
🔧 Implementer attempt —
ci-not-ready.(attempt #31, tier 2)
🔧 Implementer attempt —
blocked.Blockers:
830099ac97but dispatch base wasd8b4823b75. The implementer pushed from inside the worktree (forbidden by the git contract) OR a third party pushed during the attempt. Re-dispatch will re-prefetch and pick up the new head.830099ac97tod28778f91a🌱 Grooming: proceed — PR cleared for processing.
(check
no_duplicates, categoryno_duplicates)Scanned 253 open PRs for duplicate coverage of centralized tool version management. Anchor PR #10953 addresses a unique requirement: creating .tool-versions as single source of truth for UV, Python, Helm, and Kubeconform versions, plus CI workflow integration and test coverage. No other open PR targets this infrastructure consolidation pattern. Nearby CI improvement PRs (#10845, #10846, #10959, #10954) address execution time, security scanning, or reusable workflows — distinct from version centralization.
📋 Estimate: tier 1.
5-file multi-file change: new .tool-versions config, CI YAML restructured with a new load-versions job and updated dependencies across all existing jobs, new BDD feature scenarios, new step definitions, and a new shell helper script. Additive test work (BDD features + steps) pushes this firmly past tier 0. Reasoning complexity is low (string centralization + YAML dependency wiring), but the cross-file review surface and test fixture additions are characteristic tier-1 work. Not tier 2 — no algorithmic, concurrency, or multi-subsystem complexity. CI is green.
d28778f91ato4aa812cba9(attempt #34, tier 1)
🔧 Implementer attempt —
rebased.Pushed 1 commit:
4aa812c.✅ Approved
Reviewed at commit
4aa812c.Confidence: high.
Claimed by
merge_drive.py(pid 2329255) until2026-06-15T07:33:55.375656+00:00.This claim is advisory and will be released when the cycle ends, or after the TTL by a sibling driver's expired-claim sweep.
4aa812cba9to4ef9e17feaApproved by the controller reviewer stage (workflow 380).