The encoded large DNA can be cloned and stored in vivo, capable of write-once and stable replication for multiple retrievals, offering potential in economic data archiving. Nanopore sequencing is advantageous in data access of large DNA due to its rapidity and long-read sequencing capability. However, the data readout is commonly limited by insertion and deletion (indel) errors and sequence assembly complexity. Here, a pragmatic soft-decision data readout is presented, achieving assembly-free sequence reconstruction, indel error correction, and ultra-low coverage data readout. Specifically, the watermark is cleverly embedded within large DNA fragments, allowing for the direct localization of raw reads via watermark alignment to avoid complex read assembly. A soft-decision forward-backward algorithm is proposed, which can identify indel errors and provide probability information to the error correction code, enabling error-free data recovery. Additionally, a minimum state transition is maintained, and a read segmentation is incorporated to achieve fast information reading. The readout assays for two circular plasmids (~51Â kb) with different coding rates were demonstrated and achieved error-free recovery directly from noisy reads (error rateâ~1%) at coverage of 1-4Ã. Simulations conducted on large-scale datasets across various error rates further confirm the scalability of the method and its robust performance under extreme conditions. This readout method enables nearly single-molecule recovery of large DNA, particularly suitable for rapid readout of DNA storage.
Pragmatic soft-decision data readout of encoded large DNA.
阅读:3
作者:Ge Qi, Qin Rui, Liu Shuang, Guo Quan, Han Changcai, Chen Weigang
| 期刊: | Briefings in Bioinformatics | 影响因子: | 7.700 |
| 时间: | 2025 | 起止号: | 2025 Mar 4; 26(2):bbaf102 |
| doi: | 10.1093/bib/bbaf102 | ||
特别声明
1、本文转载旨在传播信息,不代表本网站观点,亦不对其内容的真实性承担责任。
2、其他媒体、网站或个人若从本网站转载使用,必须保留本网站注明的“来源”,并自行承担包括版权在内的相关法律责任。
3、如作者不希望本文被转载,或需洽谈转载稿费等事宜,请及时与本网站联系。
4、此外,如需投稿,也可通过邮箱info@biocloudy.com与我们取得联系。
