Preprint / Version 1

Deep conservation does not amplify proteomic detection of reverse-strand shadow ORFs in vertebrates

This article is a preprint and has not been certified by peer review.

Authors

Categories
Keywords
Shadow ORF; ncORF; Proteome; Bony fish; Conserved sequences

Abstract

Reverse-strand overlapping open reading frames ("shadow ORFs") are pervasive on the antisense strand of protein-coding genes, but whether they are translated in vertebrates has not been tested at the class level. We catalogued shadow ORFs — defined here as the longest stop-free stretch in a frame, without an initiation-codon requirement — in 318 deeply conserved human genes (orthologues required in at least four of five non-human vertebrates; 248 reach zebrafish, ~430 Myr), quantified frame openness against two null models, and ranked 1,908 six-frame candidates. To test translation without overclaiming, we ran a pre-specified mass-spectrometry enrichment test on three peptide sets of 958 each, matched on peptide length and set size: T, a high-conservation test set (27 genes, composite score 73.3–94.9); C, a low-conservation reverse control (190 genes, composite < 50); and S, a shuffled decoy set derived from T. Each set was searched under identical criteria across six public proteomes, and results were combined as the union of peptides confident in at least one dataset. T exceeded S (union OR = 2.63, P = 1.1 × 10⁻¹¹), showing that the search separates real reverse-frame sequence from shuffled sequence; this bounds the assay's statistical resolution but does not by itself establish sensitivity to genuine translation. Against C, however, T showed no class-level enrichment (union OR = 0.81, P = 0.076); with 958 peptides per set the design had 80% power to detect an odds ratio of 1.35 or greater, so smaller enrichments remain untested. The result was robust to a ≥ 9-amino-acid length restriction (OR = 0.81) and to gene-level clustering (permutation P = 0.26). Cross-dataset recurrence did not separate conserved from low-conservation candidates (10 T vs 12 C peptides confident in ≥ 3 datasets; P = 0.83), and one shuffled peptide was confident in four of six datasets, quantifying a non-trivial false-positive floor. Strand-specific total RNA-seq detected antisense transcription at 26 of the 27 T loci but at only 0.9–3.9% of sense abundance, and antisense level did not track the reproducible peptide loci. We report the recurring loci, led by EGLN1 and NFIB, as calibrated, uncertainty-annotated leads for targeted validation — leads, not proof of function.

References

Michel AM, Fox G, Kiran AM, et al. GWIPS-viz: development of a ribo-seq genome browser. Nucleic Acids Res 2014;42(D1):D859–D864. doi:10.1093/nar/gkt1035

Michel AM, Kiniry SJ, O'Connor PBF, Mullan JPA, Baranov PV. GWIPS-viz: 2018 update. Nucleic Acids Res 2018;46(D1):D823–D830. doi:10.1093/nar/gkx790

Wen B, Zhang B. PepQuery2 democratizes public MS proteomics data for rapid peptide searching. Nat Commun 2023;14(1):2213. doi:10.1038/s41467-023-37747-8

Zehentner B, Ardern Z, Kreitmeier M, Scherer S, Neuhaus K. A novel pH-regulated, unusual 603 bp overlapping protein coding gene pop is encoded antisense to ompA in Escherichia coli O157:H7 (EHEC). Front Microbiol 2020;11:377. doi:10.3389/fmicb.2020.00377

Michel AM, Choudhury KR, Firth AE, Ingolia NT, Atkins JF, Baranov PV. Observation of dually decoded regions of the human genome using ribosome profiling data. Genome Res 2012;22(11):2219–2229. doi:10.1101/gr.133249.111

Iyengar BR, Grandchamp A, Bornberg-Bauer E. How antisense transcripts can evolve to encode novel proteins. Nat Commun 2024;15(1):6187. doi:10.1038/s41467-024-50550-3

Duncan CDS, Mata J. The translational landscape of fission yeast meiosis and sporulation. Nat Struct Mol Biol 2014;21(7):641–647. doi:10.1038/nsmb.2843

Brunet MA, Brunelle M, Lucier JF, et al. OpenProt: a more comprehensive guide to explore eukaryotic coding potential and proteomes. Nucleic Acids Res 2019;47(D1):D403–D410. doi:10.1093/nar/gky936 (2021 update: Nucleic Acids Res 49(D1):D380–D388. doi:10.1093/nar/gkaa1036)

Willensdorfer M, Bürger R, Nowak MA. Phenotypic mutation rates and the abundance of abnormal proteins in yeast. PLoS Comput Biol 2007;3(11):e203. doi:10.1371/journal.pcbi.0030203

Sydow JF, Cramer P. RNA polymerase fidelity and transcriptional proofreading. Curr Opin Struct Biol 2009;19(6):732–739. doi:10.1016/j.sbi.2009.10.009

Thomas MJ, Platas AA, Hawley DK. Transcriptional fidelity and proofreading by RNA polymerase II. Cell 1998;93(4):627–637. doi:10.1016/S0092-8674(00)81191-5

Jeon C, Agarwal K. Fidelity of RNA polymerase II transcription controlled by elongation factor TFIIS. Proc Natl Acad Sci USA 1996;93(24):13677–13682. doi:10.1073/pnas.93.24.13677

Traverse CC, Ochman H. Conserved rates and patterns of transcription errors across bacterial growth states and lifestyles. Proc Natl Acad Sci USA 2016;113(12):3311–3316. doi:10.1073/pnas.1525329113

Mudge JM, Ruiz-Orera J, Prensner JR, et al. Standardized annotation of translated open reading frames. Nat Biotechnol 2022;40(7):994–1003. doi:10.1038/s41587-022-01369-0

Deutsch EW, Kok LW, Mudge JM, et al. Expanding the human proteome with microproteins and peptideins. Nature 2026;654:813–825. doi:10.1038/s41586-026-10459-x

Prensner JR, Abelin JG, Kok LW, et al. What can Ribo-seq, immunopeptidomics, and proteomics tell us about the noncanonical proteome? Mol Cell Proteomics 2023;22(9):100631. doi:10.1016/j.mcpro.2023.100631

Katoh K, Standley DM. MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol Biol Evol 2013;30(4):772–780. doi:10.1093/molbev/mst010

Frankish A, Carbonell-Sala S, Diekhans M, et al. GENCODE: reference annotation for the human and mouse genomes in 2023. Nucleic Acids Res 2023;51(D1):D942–D949. doi:10.1093/nar/gkac1071

Uhlén M, Fagerberg L, Hallström BM, et al. Tissue-based map of the human proteome. Science 2015;347(6220):1260419. doi:10.1126/science.1260419

Sabath N, Landan G, Graur D. A method for the simultaneous estimation of selection intensities in overlapping genes. PLoS ONE 2008;3(12):e3996. doi:10.1371/journal.pone.0003996

Monit C, Goldstein RA, Towers GJ, et al. Positive selection analysis of overlapping reading frames is invalid. AIDS Res Hum Retroviruses 2015;31(11):1147–1148. doi:10.1089/aid.2015.0150

Wei X, Zhang J. A simple method for estimating the strength of natural selection on overlapping genes. Genome Biol Evol 2014;6(9):2453–2461. doi:10.1093/gbe/evu294

Nelson CW, Ardern Z, Wei X. OLGenie: estimating natural selection to predict functional overlapping genes. Mol Biol Evol 2020;37(1):265–279. doi:10.1093/molbev/msaa087

Gronostajski RM. Roles of the NFI/CTF gene family in transcription and development. Gene 2000;249(1–2):31–45. doi:10.1016/s0378-1119(00)00140-2

Zhou Q, Jiang Y, Cai C, et al. Multidimensional conservation analysis decodes the expression of conserved long noncoding RNAs. Life Sci Alliance 2023;6(6):e202302002. doi:10.26508/lsa.202302002

Ruiz-Orera J, Albà MM. Conserved regions in long non-coding RNAs contain abundant translation and protein–RNA interaction signatures. NAR Genom Bioinform 2019;1(2):lqz002. doi:10.1093/nargab/lqz002

Ulitsky I. Evolution to the rescue: using comparative genomics to understand long non-coding RNAs. Nat Rev Genet 2016;17(10):601–614. doi:10.1038/nrg.2016.85

Dobin A, Davis CA, Schlesinger F, et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics 2013;29(1):15–21. doi:10.1093/bioinformatics/bts635

ENCODE Project Consortium. An integrated encyclopedia of DNA elements in the human genome. Nature 2012;489(7414):57–74. doi:10.1038/nature11247

Merino E, Balbas P, Puente JL, Bolivar F. Antisense overlapping open reading frames in genes from bacteria to human. Trends Genet 1994;10(5):160–164. doi:10.1016/0168-9525(94)90125-2

Silke J. The majority of long non-stop reading frames on the antisense strand can be explained by biased codon usage. Gene 1997;194(1):143–155. doi:10.1016/s0378-1119(97)00199-6

Rother KI, Silke J, Georgiev O, Schaffner W, Matsuo K. Influence of DNA sequence and methylation status on the open reading frames in the antisense strand of the Hsp70 and PrP genes. Biol Chem 1997;378(12):1521–1530. doi:10.1515/bchm.1997.378.12.1521

Forsdyke DR. A stem-loop kissing model for the initiation of recombination and the origin of introns. Mol Biol Evol 1995;12(5):949–958. doi:10.1007/bf00175816

Wacholder A, Carvunis AR. Biological factors and statistical limitations prevent detection of most noncanonical proteins by mass spectrometry. PLoS Biol 2023;21(8):e3002409. doi:10.1371/journal.pbio.3002409

Aggarwal S, Yadav AK. The Achilles' heel of proteogenomics: false discovery rate. Brief Bioinform 2022;23(4):bbac163. doi:10.1093/bib/bbac163

Li Q, Shortreed MR, Wenger CD, et al. A global approach for proteogenomic discovery of noncanonical reading frames and frame-shifted peptides. BMC Genomics 2016;17:1032. doi:10.1186/s12864-016-3327-5

Blakeley P, Overton IM, Hubbard SJ. Addressing statistical biases in nucleotide-derived protein databases for proteogenomic search strategies. J Proteome Res 2012;11(11):5221–5234. doi:10.1021/pr300411q

Metrics

Views: 18
Downloads: 5

Downloads

Posted

2026-08-14

How to Cite

Chen, Y., Zhao, Q., Zhang, Y., & Guo, Y. (2026). Deep conservation does not amplify proteomic detection of reverse-strand shadow ORFs in vertebrates. LangTaoSha Preprint Server. https://doi.org/10.65215/LTSpreprints.2026.08.13.000310

Download Citation

Declaration of Competing Interests

The authors declare no competing interests to disclose.