<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Resources | HENG CHEN</title><link>https://heng.ac.cn/resources/</link><atom:link href="https://heng.ac.cn/resources/index.xml" rel="self" type="application/rss+xml"/><description>Resources</description><generator>Wowchemy (https://wowchemy.com)</generator><language>en-us</language><copyright>© 2022 Oxford · HENG CHEN | Thanks for your visiting.</copyright><image><url>https://heng.ac.cn/media/icon_hu197c37e438425c83ba9767a1ce72aaf7_8389_512x512_fill_lanczos_center_3.png</url><title>Resources</title><link>https://heng.ac.cn/resources/</link></image><item><title>Enhanced TAIR12 annotation (AraBase)</title><link>https://heng.ac.cn/resources/tair12/</link><pubDate>Wed, 29 Jul 2026 00:00:00 +0000</pubDate><guid>https://heng.ac.cn/resources/tair12/</guid><description>&lt;h2 id="overview">Overview&lt;/h2>
&lt;hr>
&lt;div class=article-container>
&lt;div class=article-style>
&lt;p>The official &lt;a href="https://phoenixbioinfo.org/announcing-the-publication-of-tair12-consult-the-fully-reannotated-genome-on-the-european-nucleotide-archive/" target="_blank">TAIR12&lt;/a> genome provides (1) a completed (T2T) assembly and (2) an updated and comprehensive annotation of the &lt;em>Arabidopsis thaliana&lt;/em> genome . However, the original GFF3 file contains all annotation types (including protein-coding genes, noncoding RNAs, transposable elements, etc) in a single combined file. This structure can make it inconvenient to retrieve specific feature classes for downstream analyses. Moreover, the original protein-coding annotation did not contain explicit 5′ and 3′ UTR features. The organelle genomes (mitochondria and chloroplast) are missing in the TAIR12 files as well.&lt;p>
&lt;p>To improve accessibility and usability, the TAIR12 annotation was separated into multiple feature-specific GFF3 files. These include protein-coding genes, pseudogenes, noncoding RNAs, non-TE repeats, all transposable elements, Class I retrotransposons, Class II DNA transposons, and the major TE groups LTR, LINE, SINE, rolling-circle, and TIR elements. Separate genomic sequence files were also generated for all TEs, protein-coding genes, transcripts, CDSs, and predicted proteins.&lt;p>
&lt;p>To address the 5′ and 3′ UTR features missing issue, splice-aware UTR annotations were reconstructed for individual transcript isoforms using the exon and CDS structures in the GFF3 file. The predicted UTR intervals were then validated against the TAIR12 genome sequence and the officially distributed 5′ and 3′ UTR FASTA files. The resulting gene annotation now contains integrated five_prime_UTR and three_prime_UTR features, together with detailed validation and summary tables reporting sequence, length, chromosome, and coordinate agreement.&lt;p>
&lt;p>For the mitochondria and chloroplast chromosomes, we used the assemblies and annotations of TAIR10 version, which are completed assemblies, and merged to the TAIR12 genomes. The organelle genome annotations are in a separated file. &lt;p>
&lt;p>These processed TAIR12 resources have also been incorporated into AraBase, where users can browse and visualize genes, transcripts, UTRs, transposable elements, and other genomic features. AraBase additionally provides a convenient platform for comparing TAIR12 annotations with the widely used TAIR10/Araport11 annotation, helping users examine annotation updates, identify corresponding loci, and explore differences between genome releases.&lt;p>
&lt;p>All the scripts used in this process are provided, and the resulting files are available via &lt;a href="https://doi.org/10.6084/m9.figshare.33112340" target="_blank">Figshare&lt;/a>. The AraBase are available via &lt;a href="https://arabidopsis.ac.cn" target="_blank">arabidopsis.ac.cn&lt;/a>.&lt;p>
&lt;p>Please feel free to reach out if there are any questions.&lt;p></description></item><item><title>CapBase</title><link>https://heng.ac.cn/resources/capbase/</link><pubDate>Mon, 06 Jul 2026 00:00:00 +0100</pubDate><guid>https://heng.ac.cn/resources/capbase/</guid><description>&lt;h2 id="overview">Overview&lt;/h2>
&lt;hr>
&lt;div class=article-container>
&lt;div class=article-style>
&lt;p>We are pleased to announce the release of &lt;strong>CapBase&lt;/strong>, a developing database for &lt;em>Capsella&lt;/em> genetics, genomics, and epigenomics. The website is available at &lt;a href="https://capsella.uk/" target="_blank" rel="noopener">&lt;strong>capsella.uk&lt;/strong>&lt;/a>.&lt;/p>
&lt;p>CapBase provides a focused resource for research on &lt;em>Capsella&lt;/em>, a &lt;em>Brassicaceae&lt;/em> genus useful for studying evolution, population genomics, and epigenomics. The database currently includes resources for &lt;em>Capsella rubella&lt;/em>, &lt;em>Capsella grandiflora&lt;/em>, and &lt;em>Capsella orientalis&lt;/em>.
&lt;p>CapBase is maintained by the &lt;a href="https://www.mosherlab.com/" target="_blank">Mosher lab&lt;/a> at the &lt;a href="https://www.biology.ox.ac.uk/" target="_blank">Department of Biology, University of Oxford&lt;/a>. This database is supported by the Plant Genome Research Program of the U.S. National Science Foundation and is free for academic use.
&lt;p>For questions, comments, or collaborations, feel free to contact us by emailing emailing &lt;a href=mailto:rebecca.mosher@biology.ox.ac.uk>Prof. Rebecca Mosher&lt;/a>, or &lt;a href=mailto:heng.chen@pmb.ox.ac.uk>Heng Chen&lt;/a>, or through the &lt;a href="https://capsella.uk/contact/" target="_blank">CapBase contact page&lt;/a>.</description></item><item><title>Enhanced Araport11 annotation</title><link>https://heng.ac.cn/resources/araport11/</link><pubDate>Thu, 16 Oct 2025 00:00:00 +0000</pubDate><guid>https://heng.ac.cn/resources/araport11/</guid><description>&lt;h2 id="overview">Overview&lt;/h2>
&lt;hr>
&lt;div class=article-container>
&lt;div class=article-style>
&lt;p>The official &lt;em>Arabidopsis thaliana&lt;/em> genome annotation (&lt;a href="https://www.arabidopsis.org/download/list?dir=Genes%2FAraport11_genome_release" target="_blank">Araport11&lt;/a>) provides comprehensive information on genes, noncoding RNAs, and transposable elements (TEs). However, the distributed GFF3 file merges all genomic features, like protein-coding genes, noncoding RNAs, and transposable elements, into a single annotation file without explicit TE classification. Moreover, this lack of classification metadata makes it difficult to extract and analyze specific TE families or subclasses. Consequently, the original format is not optimal for downstream applications such as RNA-seq read alignment, TE expression analysis, or TE–gene association studies.&lt;p>
&lt;p>To address this limitation, TE classification information was incorporated into the Araport11 annotation based on the &lt;a href="https://www.arabidopsis.org/browse/transposon_families" target="_blank">TAIR Transposon Family Database&lt;/a>. The revised annotation includes a standardized RepeatMasker and EDTA style classification (classification=subclass/family/clade) tag for each TE entry. The full annotation was also separated into multiple feature-specific files (e.g., protein-coding genes, TE genes, and TEs), providing more flexibility for different genomic analyses such as gene expression quantification or TE characterization.&lt;p>
&lt;p>All the scripts used in this process are provided, and the resulting files are available via &lt;a href="https://figshare.com/articles/dataset/Enhenced_Araport11_annotation/30383395" target="_blank">Figshare&lt;/a>.&lt;p>
&lt;p>Please feel free to reach out if there are any questions.&lt;p></description></item></channel></rss>