For AI agents (and people building them)
In one sentence: this is a substrate for agents that work on Galaxy tools: call it over MCP (nine tools) or embed the library, and let it do the deterministic, verifiable parts (format, upgrade, check) while the agent does the reasoning.
Why an agent wants this
An LLM can draft a tool wrapper, but it shouldn’t guess whether the XML is valid, in canonical form, or safe at a newer profile. This project answers those deterministically and returns structured results, so the agent offloads the parts that must be correct, not plausible. The upgrade + validation framework is the most mature surface: it’s backed by per-release XSDs and a 9,373-tool evidence base, and it tells you when a change is provably safe.
Two ways in
MCP (tool calls)
Nine tools, format_tool, upgrade_tool, check_tool, convert_help_tool,
tokenize_version_tool, find_references_tool, rename_param_tool,
list_rulesets, list_rules,
take the tool XML as a string and return JSON. Nothing is written to disk. See
usage/mcp. The key signal for autonomy:
// upgrade_tool(xml=..., modernize=true) ->
{ "formatted": "<tool … profile=\"24.1\">…", "behavior_preserving": true,
"baseline_profile": "18.01", "reached_profile": "24.1",
"stopped_at": "24.1", "blocking_codes": ["24_2_fix_test_case_validation"],
"auto_fixed_codes": [], "steps_applied": [...] }
The default is minimal-bump: profile= moves only when strictly needed for
validity, and a tool that validates where it sits comes back byte-identical
(baseline_profile == reached_profile). modernize=true opts into the
behavior-preserving walk, which stops at the behaviour ceiling and never
passes the deployment ceiling (the newest profile every major public Galaxy
server runs; only an explicit target_profile exceeds it);
stopped_at/blocking_codes say where and why (each code maps to a section of
docs/profile_boundaries.md). Crossing a
behaviour boundary additionally requires allow_behavior_change=true
explicitly.
behavior_preserving (true/false/null) lets an agent decide what to
accept unattended versus surface to a human; the honest contract is in
soundness.
Library (embed it)
from galaxy_tool_refactor_registry import facade, resolve
det = facade.detect(tool_path, codes=resolve.resolve_codes(rulesets=["strict"]))
for v in det.violations:
... # v.code, v.line, v.message ; det.is_advisory(v)
Full surface and the path-vs-bytes gotcha: usage/library.
A safe-by-default loop
checkthe draft → structured findings (fixable vs advisory).format→ canonical XML (behaviour-preserving; accept freely).upgrade→ the minimal bump by default (accept freely: validity repairs plus a needed profile move only); withmodernize=true, gate onbehavior_preserving: auto-accepttrue, escalatefalse/nulland any non-emptyblocking_codes(fix the tool per the boundary reference, or ask a human before passingallow_behavior_change).Re-
checkto confirm.
The engine is deterministic and report-first, so an agent can stay conservative: prefer the proven-safe action, surface the rest.
Honest about the frontier
Agents calling the tools: shipped. That’s vision Goal 1.
Agents authoring their own rules (new codemods/checks the framework discovers and runs alongside the baked-in set): vision Goal 2: open design, not built (
galaxy-tool-refactor-mcp/docs/vision.md).upgrade’s default bumps only when validity strictly requires it; themodernizewalk stops rather than cross a breaking behaviour change it cannot prove fixed; withallow_behavior_changeit is sound for structural validity only; never treat that bump as behaviour-neutral without checkingbehavior_preserving(soundness).
Go deeper
usage/mcp · usage/library · capabilities · where this fits the ecosystem