Skip to content

[Draft] New command: "hybracter automatic" - hybrid and long samples in one go! - #131

Open
richardstoeckl wants to merge 1 commit into
gbouras13:mainfrom
richardstoeckl:main
Open

richardstoeckl wants to merge 1 commit into
gbouras13:mainfrom
richardstoeckl:main

Conversation

@richardstoeckl

@richardstoeckl richardstoeckl commented Apr 9, 2025 •

Copy link
Copy Markdown
Contributor

Hi George,

In my workflow, I sometimes have samples that have short read data available for them and sometimes there is only long read data available.
So every time I want to assemble them, I need to execute hybracter twice, instead of just letting it run over night in one go.

So I was wondering, why you separate the "hybrid" and "long-only" mode so strictly? Is this a deliberate design choice?

In this PR I wanted to present a mockup (!!) of how an "automatic" mode could work, which runs either the hybrid or the long-only pipeline, depending on the availability of short read paths in the sample.csv file (as seen in hybracter/test_data/test_hybrid_automatic_auto.csv), on a per-sample basis.

Basically, the idea is to add some logic to hybracter/workflow/rules/preflight/samples.smk, so that the SAMPLES can be split into HYBRID_SAMPLES and LONG_SAMPLES, which can then be used in wildcard_constraints and in expand() functions to choose which samples get processed by either pipeline.

Obviously, this disables some sanity checks regarding the required existence of paths/files, but maybe one could communicate this mode to be best suited for experienced people?

For this PR, I just did a very crude mockup, as I did not want to waste too much time if this Idea is not something you want to pursue further.
Please excuse the horrible python code in samplesFromCsvAutomatic(), parseSamples(), and filter_samples_by_workflow_type() and the copy-and-pasted-together rules and snaketool code, I tried my best to understand the inner workings.
The mockup works only with hybracter automatic -i hybracter/test_data/test_hybrid_automatic_auto.csv --no_medaka --databases /path/to/databases/ --auto, again, to not waste too much time. It "works" in the sense that it runs through on my machine without errors, but I have not checked for mistakes or on other machines.

If you want to pursue this idea further, let me know, and I can polish everything up a bit more.

Best wishes,
Richard

PS: "hybracter automatic" is probably not the best name, especially with regards to confusion about the --auto parameter, but I couldn't think of something better for now.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant