Is it strictly necessary to filter out taxa failing the symtest if the topology has biological sense? #529
|
Hi everyone! I am running a phylogenetic analysis on a SNP dataset (IQ-TREE 3). When running On the other hand, when I build the Maximum Likelihood tree keeping all the samples (including those that failed the symmetry test), the resulting topology makes biological sense. So, in SNP datasets, is it strictly necessary to prune taxa that fail the symtest if the tree topology is robust and biologically meaningful? What are the risks of ignoring these failures vs. the risks of reducing taxon sampling by continuously removing samples? Thank you for your time! DJG |
Replies: 2 comments
|
Hi David, I think you've asked the million dollar question there! To answer your titular question directly - I think the answer is 'no'. To me, what it comes down to is answering the question "When is a error- or bias-prone estimate of a topology better than no estimate?". And bear in mind that every topology is error prone, it's just that usually: (i) nobody knows how error prone; and (ii) nobody checks anyway. With respect to your questions, I wouldn't put much stock in 'robust', if by robust you mean 'has high bootstrap values'. Any dataset of reasonable size should have high bootstrap values. But biologically meaningful is a good sanity check. My own feeling is that the best you can do is to measure these kinds of aspects of bias, and if you end up in your situation where it's a choice between no estimate and a potentially biased one, go ahead with the biased one but keep the potential bias in mind. For example, if you map the base composition onto the tree, and see that taxa with extreme base composition are being lumped together, then you have to keep in mind that you can't really know (with stationary models), if that's the biology or the bias doing the talking. A final point - we have a follow up to the symtest paper which might be of interest: https://academic.oup.com/sysbio/advance-article/doi/10.1093/sysbio/syag010/8465386 This shows that when we simulate data, the only time the topology seems to go really bad are the cases where there's been quite dramatic correlated shifts in base composition in different parts of the tree. I hope some of that helps, Rob |
|
Hi, Thank you so much for your response! It’s great to know that maintaining taxon sampling is the preferred path here. I will proceed with the full dataset, keep an eye out for any artificial clustering, and move forward with the analysis. Because, as you said (and I think so too), it's better to have a potentially biased phylogeny than not having one. I will check out your follow-up paper on the symtest as well. Thanks again for your time and guidance! Best regards, |
Hi David,
I think you've asked the million dollar question there! To answer your titular question directly - I think the answer is 'no'.
To me, what it comes down to is answering the question "When is a error- or bias-prone estimate of a topology better than no estimate?". And bear in mind that every topology is error prone, it's just that usually: (i) nobody knows how error prone; and (ii) nobody checks anyway.
With respect to your questions, I wouldn't put much stock in 'robust', if by robust you mean 'has high bootstrap values'. Any dataset of reasonable size should have high bootstrap values. But biologically meaningful is a good sanity check.
My own feeling is that the best you can d…