Improving speed and efficiency is a constant goal. We've redesigned dependencies to be faster, we've built ReDos detection pipelines, and we've built timing sensors into every sub-phase of the engine to explore any abnormality that might pop up. While this initiative started to ensure that this scan could be in the CI pipeline, we found the aggregate data to be a gold mine.
We've documented the phases enough now, that speed helps us find any miswiring early. Our engine's speed follows a linear log-log relationship between repo LOC (of any language) versus scan time. We can scan any repo, sort by scanning time for every metric and systematically address the outliers. This allows us to detect issues on brand new never scanned repos (..these files scanned 10x slower, likely contains redos errors which correlate with extraction errors) so now speed, and the deviation from it, is also a metric to flag new to the system issues.
Overall, we've used this system to create hypotheses around speed, test them out at scale and determine how helpful they would be or not, we've deferred some ideas that warranted an assessment and left them up for others to check. Who knows, maybe they will become more of a bottleneck as capabilities grow.
Issues
- 1
- 3
- 3
- 1
- 1