















Abstract:As DNA data storage advances toward practical deployment, minimizing sequencing coverage depth is critical for reducing operational costs and retrieval latency. We study the random access problem of recovering a specific information strand from a DNA-based storage system. In this setting, $k$ information strands are encoded into $n$ strands using a generator matrix $G$, and each sequencing read returns one encoded strand sampled uniformly at random with replacement. We derive an exact formula for the expected number of samples required to recover a specific information strand, yielding an $O(n)$-time algorithm for fixed field size $q$ and dimension $k$. We further obtain explicit formulas for the average and maximum expected number of samples, enabling an efficient search for optimal generator matrices for small parameters. We present new constructions that improve the best-known upper bounds from $0.8815k$ to $0.8811k$ for $k=3$, and from $0.8637k$ to $0.8629k$ for $k=4$, for sufficiently large $q$. We also establish a tighter lower bound on the expected number of samples, which in particular proves the optimality of the simple parity code when $n=k+1$ over any field size $q$. Finally, for the non-random access setting, we derive new lower bounds and constructions that characterize the asymptotic behavior of the expected number of samples required to recover all information strands.
From: Chen Wang [view email]
[v1]
Sun, 11 Jan 2026 20:12:08 UTC (19 KB)
[v2]
Wed, 12 Aug 2026 14:43:09 UTC (33 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。