[Draft] New command: "hybracter automatic" - hybrid and long samples in one go! - #131
Open
richardstoeckl wants to merge 1 commit into
Open
richardstoeckl wants to merge 1 commit into
richardstoeckl wants to merge 1 commit into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Hi George,
In my workflow, I sometimes have samples that have short read data available for them and sometimes there is only long read data available.
So every time I want to assemble them, I need to execute hybracter twice, instead of just letting it run over night in one go.
So I was wondering, why you separate the "hybrid" and "long-only" mode so strictly? Is this a deliberate design choice?
In this PR I wanted to present a mockup (!!) of how an "automatic" mode could work, which runs either the hybrid or the long-only pipeline, depending on the availability of short read paths in the sample.csv file (as seen in hybracter/test_data/test_hybrid_automatic_auto.csv), on a per-sample basis.
Basically, the idea is to add some logic to hybracter/workflow/rules/preflight/samples.smk, so that the
SAMPLEScan be split intoHYBRID_SAMPLESandLONG_SAMPLES, which can then be used inwildcard_constraintsand inexpand()functions to choose which samples get processed by either pipeline.Obviously, this disables some sanity checks regarding the required existence of paths/files, but maybe one could communicate this mode to be best suited for experienced people?
For this PR, I just did a very crude mockup, as I did not want to waste too much time if this Idea is not something you want to pursue further.
Please excuse the horrible python code in
samplesFromCsvAutomatic(),parseSamples(), andfilter_samples_by_workflow_type()and the copy-and-pasted-together rules and snaketool code, I tried my best to understand the inner workings.The mockup works only with
hybracter automatic -i hybracter/test_data/test_hybrid_automatic_auto.csv --no_medaka --databases /path/to/databases/ --auto, again, to not waste too much time. It "works" in the sense that it runs through on my machine without errors, but I have not checked for mistakes or on other machines.If you want to pursue this idea further, let me know, and I can polish everything up a bit more.
Best wishes,
Richard
PS: "hybracter automatic" is probably not the best name, especially with regards to confusion about the
--autoparameter, but I couldn't think of something better for now.