Milestones
List view
The engine's architecture roadmap: every system relation as a named fact channel (#3249), a channel registry so new channels don't edit shared seams (#3383), and what the facts unlock -- true dead-code detection (#3266) and nativeness scoring of generated code (#3123). The COBOL-to-Java generator is built on this tree.
No due date•19/23 issues closedDo GitGalaxy's risk signals predict real defects and incidents, or are they just numbers? Validation against independently identified events over Git history (epic #2982), outcome-label pilots, calibration contracts that admit a signal only on demonstrated defect-lift beyond size, and cleaning out dead or pinned metrics.
No due date•43/50 issues closedBroadening the GlassWorm-style supply-chain worm detector (epic #1171): what a worm is after (credential files), how it spreads and survives (self-propagation, persistence), how it exfiltrates (secret dumps), across more languages and harder to dodge.
No due date•2/8 issues closedThe adversarial estates built to break a COBOL-to-Java port, and the hard cases they expose: CICS transactions (#3989), i18n and code pages (#3988), bidirectional 3270 screens, collation, reserved words, units the engine misses. If a port survives these, it survives your estate.
No due date•15/19 issues closedCan the porting loop prove code it has never seen, on the first try? Fresh public estates through the trial protocol (first-try numbers frozen before any porting), test inputs generated from record layouts, and the attempts-to-proof histogram. Epics #3755 and #3803.
No due date•6/10 issues closedOther tools let a model rewrite COBOL and claim it works. Here every port comes with its evidence: a byte-for-byte replay against the original, coverage, a mutation score that measures how much the proof really checks, and a person's sign-off. The Java is only half the product; the proof is the other half. Epic #4055.
No due date•8/17 issues closedTraditional AST parsers require mathematically perfect grammar and syntax. We don't even require source code. We extract structure and language keywords from strings of text, doesn't matter if the pattern is in English or binary. It remains to be assessed how much extraction we can get though, we envision some information loss.
No due date•0/3 issues closedGitGalaxy reads IBM z/OS COBOL, JCL and CICS. Real estates are older and wider than that. This milestone teaches the engine the rest of the back catalogue: - **Other vendors' COBOL:** Fujitsu NetCOBOL, Unisys (OS 2200 / MCP, DMSII), HPE NonStop (Tandem SCOBOL, Enscribe) - **Other ecosystems:** IBM i (RPG, CL, DDS), Software AG (Natural, Adabas) - **Pre-relational databases:** CA IDMS, Datacom - **Report writers and 4GLs:** Easytrieve, FOCUS, Mark IV / Vision:Builder - **Fixed-format and utilities:** Fortran 77, DFSORT / SyncSort control-card lineage Column positions still matter here, and that's where a positional reader beats a grammar. **Belongs here if:** the engine doesn't read a legacy language, dialect or data store, or reads it wrong. If the code *is* read and the Java port breaks, that's **Nightmare Mode**. If the proof is weak, that's **Run Verified**. If there's no source at all, that's **Achievement Unlocked: Source Code Optional**.
No due date•0/10 issues closedDefects in GitGalaxy's function call graph found by comparing it against a compiler or type-checker reference: pyan3 for Python, the TypeScript compiler's type checker (tests/tools/ts_callgraph.js) for TypeScript, and future references for other languages. The comparison runs as the Level 2 call-resolution gate (tests/tools/call_graph_resolution.py, graph-accuracy-audit.yml). Every issue it surfaces goes here: a wrong confident link, a link into a bodyless signature, a call the engine never extracts, a reference-tool gap. Also here are the PRs that add a language's reference or fix what it finds. Numbers quoted in these issues are against a reference, not ground truth. Check disagreements against source before claiming them (CLAUDE.md, "Comparative-correctness claims").
No due date•64/80 issues closedImproving speed and efficiency is a constant goal. We've redesigned dependencies to be faster, we've built ReDos detection pipelines, and we've built timing sensors into every sub-phase of the engine to explore any abnormality that might pop up. While this initiative started to ensure that this scan could be in the CI pipeline, we found the aggregate data to be a gold mine. We've documented the phases enough now, that speed helps us find any miswiring early. Our engine's speed follows a linear log-log relationship between repo LOC (of any language) versus scan time. We can scan any repo, sort by scanning time for every metric and systematically address the outliers. This allows us to detect issues on brand new never scanned repos (..these files scanned 10x slower, likely contains redos errors which correlate with extraction errors) so now speed, and the deviation from it, is also a metric to flag new to the system issues. Overall, we've used this system to create hypotheses around speed, test them out at scale and determine how helpful they would be or not, we've deferred some ideas that warranted an assessment and left them up for others to check. Who knows, maybe they will become more of a bottleneck as capabilities grow.
No due date•22/28 issues closedWorthwhile work off the main questlines -- self-contained, most of it a good way into the codebase: - **5-Tier Knowledge Graph Hierarchy (epic #108):** `SignalProcessor` flattens Function -> Class -> File -> Folder -> Repo into a File-level score, skipping the Class level. #109 the risk-aggregation math, #110 class-level risk vectors in the SQLite/output recorders, #111 the docs. - **ML-Ops positive controls (#96):** a defanged malware sample library in CI, so a structural-signature change can never silently drift the threat classifier. - **Scanning / UX:** a full scan mode ignoring the default vendor/generated exclusions (#136), a Sankey diagram of what was scanned (#137), a warning for uncommitted files outside git-based scanning (#138). - **Small fixes and cleanup:** Ruff trailing whitespace (#499), an orphaned README trigger scanner (#513), PowerShell functions missed after the aperture change (#3303), a tiktoken import failure that takes the engine down (#3791), raising the minimum Python to 3.11 (#3322), the README's strategy narrative (#3255).
No due date•30/43 issues closedEvolving GitGalaxy from a structural static analysis tool into a compliance-ready, air-gap-native DevSecOps platform for highly regulated sectors (Aerospace, Defense, FinTech, MedTech) -- tracked under epic #75. Traditional AST-based analysis can't run against decades-old, uncompilable defense/aerospace codebases. This milestone extracts Structural Signatures for legacy and modern safety-critical languages (no working toolchain required) and layers a domain-specific risk ontology on top: - **Phase 1 -- Defense Primitives (language dictionaries):** JOVIAL/CMS-2 (#77), VHDL (#78), Lustre/SCADE (#81), HAL/S (#82), CORAL 66 (#1141), PL/I (#1142), ATLAS (#1143). Ada/SPARK (#76) and AGC Assembly already shipped. - **Phase 2 -- Domain-Specific Ontology Auto-Binning Engine (#84):** classify a file's architectural role (flight software, ground control, avionics bus, ...) instead of treating every repository as generic source code. - **Phase 3 -- MISRA C/C++ Structural Signatures & Liability Floor (#80):** a hard compliance gate for the automotive/aerospace C/C++ safety-critical coding standard. SARIF export, SBOM generation, and CI/CD delta-gating (Phases 4-5) already shipped under v2.4.0 -- see #75's own status note for the current remaining scope. **Note:** #1142's PL/I signatures overlap with #2502 (Legacy Modernization, CICS-driven) -- see #2516 for the open consolidation question before either ships.
No due date•3/12 issues closed