Replies: 17 comments 9 replies
|
Interesting! In my python monorepo we use Regarding some of the tooling you mention above, I am thinking that a great first start for some use cases could maybe be to get |
|
I haven't heard of polylith before, it looks 😍 Just after I discovered ruff analyze graph looks promising. |
|
Here is another one, only a couple of months old: 🔍 CFG-based dead code detection – Find unreachable code after exhaustive if-elif-else chains |
|
Yes, thank you for opening the discussion! I don't have much to add at the moment, but I'm excited to follow the thread and take a look at the tools mentioned 😄 |
|
https://github.com/allenanswerzq/llmcc |
|
https://github.com/NathanDuboisset/importee |
|
https://github.com/matzehuels/stacktower Inspired by XKCD #2347, Stacktower renders dependency graphs as physical towers where blocks rest on what they depend on. Your application sits at the top, supported by libraries below—all the way down to that one critical package maintained by some dude in Nebraska. |
|
Related for JS/TS (written in Rust) Fallow turns a JS/TS repository into a trusted quality report: health score, changed-code risk, hotspots, duplication, architecture issues, dependency hygiene, and cleanup opportunities. It helps you answer: What changed? Fallow is built for maintainers, CI pipelines, editors, and AI agents that need structured evidence instead of guesses. No AI inside the analyzer. Fallow produces deterministic findings, typed output contracts, and traceable explanations that downstream tools can trust. |
|
Deply is a static code analysis tool for Python that helps you communicate, visualize and enforce architectural decisions in your projects. You can freely define your architectural layers over classes and which rules should apply to them. For example, you can use Deply to ensure that modules/packages in your project are truly independent of each other to make them easier to reuse. Deply can be used in a CI pipeline to make sure a pull request does not violate any of the architectural rules you defined. With the optional Mermaid formatter you can visualize your layers, rules and violations. by @Vashkatsi |
|
FlakyDetector is a research-oriented tool designed to identify non-deterministic (flaky) tests in Python codebases. Instead of relying on historical CI execution data, the analyzer parses the Abstract Syntax Tree (AST) of the source code, extracts scientific features, and classifies them using a CatBoost ML model. Unlike ordinary linters, FlakyDetector hunts for architectural anti-patterns: race conditions, resource leaks, global state dependencies, and high cyclomatic complexity (Test Smells). |
|
Most linters catch bad code inside files. Fensu catches architectural drift: code crossing the wrong boundary, living in the wrong module, or growing into the wrong shape. As a repository grows, code moves, lessons get forgotten, and the mental map decays. Tests preserve behavior and types preserve interfaces. Fensu makes the repository's architectural expectations executable. Fensu enforces:
It ships a coherent default architecture rather than a blank rule framework, then lets projects disable, extend, or replace parts deliberately. |
|
https://github.com/repowise-dev/repowise by @RaghavChamadiya Every file is scored 1-10 from 25 deterministic markers (McCabe complexity, brain methods, LCOM4 cohesion, god classes, native Rabin-Karp clone detection, untested hotspots, change entropy, prior-defect history and more), split into three lenses: defect risk, maintainability, and performance (static N+1 and I/O-in-loop risk traced across files through the call graph, where file-local linters found 0 of the cross-function cases repowise surfaced 557 of). |
|
Correct AI-generated code can still accumulate duplicated abstractions and structural decay that tests miss. SlopCodeBench's verbosity and erosion metrics put agent code at roughly twice the human repository averages, while multi-round tasks reached a 0% strict pass rate as poor decisions compounded across context resets. |
|
https://github.com/officefloor/ImpactGate by @sagenschneider |
Uh oh!
There was an error while loading. Please reload this page.
The landscape of Python software quality tooling is currently defined by two contrasting forces: high-velocity convergence and deep specialization. The recent, rapid adoption of Ruff has solved the long-standing community problem of coordinating dozens of separate linters and formatters, establishing a unified, high-performance axis for standard code quality.
A second category of tools continues to operate in necessary, but isolated, silos. Tools dedicated to architectural enforcement and deep structural metrics, such as:
These projects address fundamental challenges of code maintainability, evolvability, and architectural debt that extend beyond the scope of fast, stylistic linting. The success of Ruff now presents the opportunity to foster a cross-tool discussion focused not just on syntax, but on structure.
Specialized quality tools are vital for long-term maintainability and risk assessment. Tools like
import-linterandtachmitigate technical risk by enforcing architectural rules, preventing systemic decay, and reducing change costs. Complexity and cohesion metrics from tools such ascomplexipy,lcom, andcohesionquantitatively flag overly complex or highly coupled components, acting as early warning systems for technical debt. By analysing the combined outputs, risk assessment shifts to predictive modelling: integrating data from individual tools (e.g.,import-linterviolations,complexipyscores) creates a multi-dimensional risk score. Overlaying these results, such as identifying modules that are both low in cohesion and involved intach-flagged dependency cycles, generates a "heat map" of technical debt. This unified approach, empirically validated against historical project data like bug frequency and commit rates can yield a predictive risk assessment. It identifies modules that are not just theoretically complex but empirically confirmed sources of instability, transforming abstract quality metrics into concrete, prioritized refactoring tasks for the riskiest codebase components.Reasons to Connect
Bring the maintainers and core users of these diverse tools into a shared discussion.
Increasing Tool Visibility and Sustainability: Specialized tools often rely on small, dedicated contributor pools and suffer from knowledge isolation, confining technical debate to their specific GitHub repository. A broader discussion provides these projects with critical outreach, exposure to a wider user base, and a stronger pipeline of new contributors, ensuring their long-term sustainability.
Let's start the conversation on how to 'measure' maintainable, and architecturally sound Python code.
And keep Goodhart's law: "When a measure becomes a target, it ceases to be a good measure" in mind ;-)
All reactions