<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:media="http://search.yahoo.com/mrss/" xmlns:ynews="http://news.yahoo.com/rss/">
    <channel>
        <title>Nova Reader - Subject</title>
        <link>https://www.novareader.co</link>
        <description>Default RSS Feed</description>
        <language>en-us</language>
        <copyright>Newgen KnowledgeWorks</copyright>
        <item>
            <title><![CDATA[StreptomeDB 3.0: an updated compendium of streptomycetes natural products]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765741441869-e74f2f81-6379-4439-a9d5-51507910a1a4/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa868</link>
            <description><![CDATA[<p class="para" id="N65541">Antimicrobial resistance is an emerging global health threat necessitating the rapid development of novel antimicrobials. Remarkably, the vast majority of currently available antibiotics are natural products (NPs) isolated from streptomycetes, soil-dwelling bacteria of the genus <i>Streptomyces</i>. However, there is still a huge reservoir of streptomycetes NPs which remains pharmaceutically untapped and a compendium thereof could serve as a source of inspiration for the rational design of novel antibiotics. Initially released in 2012, StreptomeDB (http://www.pharmbioinf.uni-freiburg.de/streptomedb) is the first and only public online database that enables the interactive phylogenetic exploration of streptomycetes and their isolated or mutasynthesized NPs. In this third release, there are substantial improvements over its forerunners, especially in terms of data content. For instance, about 2500 unique NPs were newly annotated through manual curation of about 1300 PubMed-indexed articles, published in the last five years since the second release. To increase interoperability, StreptomeDB entries were hyperlinked to several spectral, (bio)chemical and chemical vendor databases, and also to a genome-based NP prediction server. Moreover, predicted pharmacokinetic and toxicity profiles were added. Lastly, some recent real-world use cases of StreptomeDB are highlighted, to illustrate its applicability in life sciences.</p>]]></description>
            <pubDate><![CDATA[2020-10-13T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[PheLiGe: an interactive database of billions of human genotype–phenotype associations]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765741398253-35ef7a0c-fce0-4239-a619-4e8fadaf2710/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1086</link>
            <description><![CDATA[<p class="para" id="N65541">Genome-wide association studies have provided a vast array of publicly available SNP × phenotype association results. However, they are often in disparate repositories and formats, making downstream analyses difficult and time consuming. PheLiGe (https://phelige.com) is a database that provides easy access to such results via a web interface. The underlying database currently stores &gt;75 billion genotype–phenotype associations from 7347 genome-wide and 1.2 million region-wide (e.g. <i>cis</i>-eQTL) association scans. The web interface allows for investigation of regional genotype-phenotype associations across many phenotypes, giving insights into the biological function affected by the variant in question. Furthermore, PheLiGe can compare regional patterns of association between different traits. This analysis can ascertain whether a co-association is due to pleiotropy or linkage. Moreover, comparison of association patterns for a complex trait of interest and gene expression and protein levels can implicate causal genes.</p>]]></description>
            <pubDate><![CDATA[2020-11-27T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[BRENDA, the ELIXIR core data resource in 2021: new developments and updates]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765741355937-2dcba8cd-bf11-487b-8215-1803c3dc019f/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1025</link>
            <description><![CDATA[<p class="para" id="N65541">The BRENDA enzyme database (https://www.brenda-enzymes.org), established in 1987, has evolved into the main collection of functional enzyme and metabolism data. In 2018, BRENDA was selected as an ELIXIR Core Data Resource. BRENDA provides reliable data, continuous curation and updates of classified enzymes, and the integration of newly discovered enzymes. The main part contains &gt;5 million data for ∼90 000 enzymes from ∼13 000 organisms, manually extracted from ∼157 000 primary literature references, combined with information of text and data mining, data integration, and prediction algorithms. Supplements comprise disease-related data, protein sequences, 3D structures, genome annotations, ligand information, taxonomic, bibliographic, and kinetic data. BRENDA offers an easy access to enzyme information from quick to advanced searches, text- and structured-based queries for enzyme-ligand interactions, word maps, and visualization of enzyme data. The BRENDA Pathway Maps are completely revised and updated for an enhanced interactive and intuitive usability. The new design of the Enzyme Summary Page provides an improved access to each individual enzyme. A new protein structure 3D viewer was integrated. The prediction of the intracellular localization of eukaryotic enzymes has been implemented. The new EnzymeDetector combines BRENDA enzyme annotations with protein and genome databases for the detection of eukaryotic and prokaryotic enzymes.</p>]]></description>
            <pubDate><![CDATA[2020-11-19T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Mouse Genome Database (MGD): Knowledgebase for mouse–human comparative biology]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765741326283-88bd7f1b-f6b7-47e9-94b0-9dbf0674b832/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1083</link>
            <description><![CDATA[<p class="para" id="N65541">The Mouse Genome Database (MGD; http://www.informatics.jax.org) is the community model organism knowledgebase for the laboratory mouse, a widely used animal model for comparative studies of the genetic and genomic basis for human health and disease. MGD is the authoritative source for biological reference data related to mouse genes, gene functions, phenotypes and mouse models of human disease. MGD is the primary source for official gene, allele, and mouse strain nomenclature based on the guidelines set by the International Committee on Standardized Nomenclature for Mice. MGD’s biocuration scientists curate information from the biomedical literature and from large and small datasets contributed directly by investigators. In this report we describe significant enhancements to the content and interfaces at MGD, including (i) improvements in the Multi Genome Viewer for exploring the genomes of multiple mouse strains, (ii) inclusion of many more mouse strains and new mouse strain pages with extended query options and (iii) integration of extensive data about mouse strain variants. We also describe improvements to the efficiency of literature curation processes and the implementation of an information portal focused on mouse models and genes for the study of COVID-19.</p>]]></description>
            <pubDate><![CDATA[2020-11-24T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The international glycan repository GlyTouCan version 3.0]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765741158667-c6c9c8bf-8612-41bc-99a4-1a9d9f56d704/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa947</link>
            <description><![CDATA[<p class="para" id="N65541">Glycans serve important roles in signaling events and cell-cell communication, and they are recognized by lectins, viruses and bacteria, playing a variety of roles in many biological processes. However, there was no system to organize the plethora of glycan-related data in the literature. Thus GlyTouCan (https://glytoucan.org) was developed as the international glycan repository, allowing researchers to assign accession numbers to glycans. This also aided in the integration of glycan data across various databases. GlyTouCan assigns accession numbers to glycans which are defined as sets of monosaccharides, which may or may not be characterized with linkage information. GlyTouCan was developed to be able to recognize any level of ambiguity in glycans and uniquely assign accession numbers to each of them, regardless of the input text format. In this manuscript, we describe the latest update to GlyTouCan in version 3.0, its usage, and plans for future development.</p>]]></description>
            <pubDate><![CDATA[2020-10-30T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[OrthoDB in 2020: evolutionary and functional annotations of orthologs]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765741014290-f30fa868-b957-4f69-a112-88c681c0ec4f/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1009</link>
            <description><![CDATA[<p class="para" id="N65541">OrthoDB provides evolutionary and functional annotations of orthologs, inferred for a vast number of available organisms. OrthoDB is leading in the coverage and genomic diversity sampling of Eukaryotes, Prokaryotes and Viruses, and the sampling of Bacteria is further set to increase three-fold. The user interface has been enhanced in response to the massive growth in data. OrthoDB provides three views on the data: (i) a list of orthologous groups related to a user query, which are now arranged to visualize their hierarchical relations, (ii) a detailed view of an orthologous group, now featuring a Sankey diagram to facilitate navigation between the levels of orthology, from more finely-resolved to more general groups of orthologs, as well as an arrangement of orthologs into an interactive organism taxonomy structure, and (iii) we added a gene-centric view, showing the gene functional annotations and the pair-wise orthologs in example species. The OrthoDB standalone software for delineation of orthologs, Orthologer, is freely available. Online BUSCO assessments and mapping to OrthoDB of user-uploaded data enable interactive exploration of related annotations and generation of comparative charts. OrthoDB strives to predict orthologs from the broadest coverage of species, as well as to extensively collate available functional annotations, and to compute evolutionary annotations such as evolutionary rate and phyletic profile. OrthoDB data can be assessed via SPARQL RDF, REST API, downloaded or browsed online from https://orthodb.org.</p>]]></description>
            <pubDate><![CDATA[2020-11-16T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[CMNPD: a comprehensive marine natural products database towards facilitating drug discovery from the ocean]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765741009145-07110853-45fa-4385-8eda-25287e12c5e5/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa763</link>
            <description><![CDATA[<p class="para" id="N65541">Marine organisms are expected to be an important source of inspiration for drug discovery after terrestrial plants and microorganisms. Despite the remarkable progress in the field of marine natural products (MNPs) chemistry, there are only a few open access databases dedicated to MNPs research. To meet the growing demand for mining and sharing for MNPs-related data resources, we developed CMNPD, a comprehensive marine natural products database based on manually curated data. CMNPD currently contains more than 31 000 chemical entities with various physicochemical and pharmacokinetic properties, standardized biological activity data, systematic taxonomy and geographical distribution of source organisms, and detailed literature citations. It is an integrated platform for structure dereplication (assessment of novelty) of (marine) natural products, discovery of lead compounds, data mining of structure-activity relationships and investigation of chemical ecology. Access is available through a user-friendly web interface at https://www.cmnpd.org. We are committed to providing a free data sharing platform for not only professional MNPs researchers but also the broader scientific community to facilitate drug discovery from the ocean.</p>]]></description>
            <pubDate><![CDATA[2020-09-28T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[DrugCentral 2021 supports drug discovery and repositioning]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740831987-f69848e8-b27b-4710-8fca-333f6a06f6e5/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa997</link>
            <description><![CDATA[<p class="para" id="N65541">DrugCentral is a public resource (http://drugcentral.org) that serves the scientific community by providing up-to-date drug information, as described in previous papers. The current release includes 109 newly approved (October 2018 through March 2020) active pharmaceutical ingredients in the US, Europe, Japan and other countries; and two molecular entities (e.g. mefuparib) of interest for COVID19. New additions include a set of pharmacokinetic properties for ∼1000 drugs, and a sex-based separation of side effects, processed from FAERS (FDA Adverse Event Reporting System); as well as a drug repositioning prioritization scheme based on the market availability and intellectual property rights forFDA approved drugs. In the context of the COVID19 pandemic, we also incorporated REDIAL-2020, a machine learning platform that estimates anti-SARS-CoV-2 activities, as well as the ‘drugs in news’ feature offers a brief enumeration of the most interesting drugs at the present moment. The full database dump and data files are available for download from the DrugCentral web portal.</p>]]></description>
            <pubDate><![CDATA[2020-11-05T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[RASP: an atlas of transcriptome-wide RNA secondary structure probing data]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740825666-09dc18d0-15bd-4b23-9de3-4b2843208e97/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa880</link>
            <description><![CDATA[<p class="para" id="N65541">RNA molecules fold into complex structures that are important across many biological processes. Recent technological developments have enabled transcriptome-wide probing of RNA secondary structure using nucleases and chemical modifiers. These approaches have been widely applied to capture RNA secondary structure in many studies, but gathering and presenting such data from very different technologies in a comprehensive and accessible way has been challenging. Existing RNA structure probing databases usually focus on low-throughput or very specific datasets. Here, we present a comprehensive RNA structure probing database called RASP (<b>R</b>NA <b>A</b>tlas of <b>S</b>tructure <b>P</b>robing) by collecting 161 deduplicated transcriptome-wide RNA secondary structure probing datasets from 38 papers. RASP covers 18 species across animals, plants, bacteria, fungi, and also viruses, and categorizes 18 experimental methods including DMS-seq, SHAPE-Seq, SHAPE-MaP, and icSHAPE, etc. Specially, RASP curates the up-to-date datasets of several RNA secondary structure probing studies for the RNA genome of SARS-CoV-2, the RNA virus that caused the on-going COVID-19 pandemic. RASP also provides a user-friendly interface to query, browse, and visualize RNA structure profiles, offering a shortcut to accessing RNA secondary structures grounded in experimental data. The database is freely available at http://rasp.zhanglab.net.</p>]]></description>
            <pubDate><![CDATA[2020-10-17T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[GreenPhylDB v5: a comparative pangenomic database for plant genomes]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740767946-1711a242-b0e0-464e-87ea-57274ea22542/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1068</link>
            <description><![CDATA[<p class="para" id="N65541">Comparative genomics is the analysis of genomic relationships among different species and serves as a significant base for evolutionary and functional genomic studies. GreenPhylDB (https://www.greenphyl.org) is a database designed to facilitate the exploration of gene families and homologous relationships among plant genomes, including staple crops critically important for global food security. GreenPhylDB is available since 2007, after the release of the <i>Arabidopsis thaliana</i> and <i>Oryza sativa</i> genomes and has undergone multiple releases. With the number of plant genomes currently available, it becomes challenging to select a single reference for comparative genomics studies but there is still a lack of databases taking advantage several genomes by species for orthology detection. GreenPhylDBv5 introduces the concept of comparative pangenomics by harnessing multiple genome sequences by species. We created 19 pangenes and processed them with other species still relying on one genome. In total, 46 plant species were considered to build gene families and predict their homologous relationships through phylogenetic-based analyses. In addition, since the previous publication, we rejuvenated the website and included a new set of original tools including protein-domain combination, tree topologies searches and a section for users to store their own results in order to support community curation efforts.</p>]]></description>
            <pubDate><![CDATA[2020-11-25T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Rfam 14: expanded coverage of metagenomic, viral and microRNA families]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740700991-cc799198-6c00-4d2d-8fa9-72eeaed5ee56/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1047</link>
            <description><![CDATA[<p class="para" id="N65541">Rfam is a database of RNA families where each of the 3444 families is represented by a multiple sequence alignment of known RNA sequences and a covariance model that can be used to search for additional members of the family. Recent developments have involved expert collaborations to improve the quality and coverage of Rfam data, focusing on microRNAs, viral and bacterial RNAs. We have completed the first phase of synchronising microRNA families in Rfam and miRBase, creating 356 new Rfam families and updating 40. We established a procedure for comprehensive annotation of viral RNA families starting with <i>Flavivirus</i> and <i>Coronaviridae</i> RNAs. We have also increased the coverage of bacterial and metagenome-based RNA families from the ZWD database. These developments have enabled a significant growth of the database, with the addition of 759 new families in Rfam 14. To facilitate further community contribution to Rfam, expert users are now able to build and submit new families using the newly developed Rfam Cloud family curation system. New Rfam website features include a new sequence similarity search powered by RNAcentral, as well as search and visualisation of families with pseudoknots. Rfam is freely available at https://rfam.org.</p>]]></description>
            <pubDate><![CDATA[2020-11-19T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[DDBJ update: streamlining submission and access of human data]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740611856-06aeaf7d-f64f-4bab-898f-8a6fe8e16931/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa982</link>
            <description><![CDATA[<p class="para" id="N65541">The Bioinformation and DDBJ Center (DDBJ Center, https://www.ddbj.nig.ac.jp) provides databases that capture, preserve and disseminate diverse biological data to support research in the life sciences. This center collects nucleotide sequences with annotations, raw sequencing data, and alignment information from high-throughput sequencing platforms, and study and sample information, in collaboration with the National Center for Biotechnology Information (NCBI) and the European Bioinformatics Institute (EBI). This collaborative framework is known as the International Nucleotide Sequence Database Collaboration (INSDC). In collaboration with the National Bioscience Database Center (NBDC), the DDBJ Center also provides a controlled-access database, the Japanese Genotype–phenotype Archive (JGA), which archives and distributes human genotype and phenotype data, requiring authorized access. The NBDC formulates guidelines and policies for sharing human data and reviews data submission and use applications. To streamline all of the processes at NBDC and JGA, we have integrated the two systems by introducing a unified login platform with a group structure in September 2020. In addition to the public databases, the DDBJ Center provides a computer resource, the NIG supercomputer, for domestic researchers to analyze large-scale genomic data. This report describes updates to the services of the DDBJ Center, focusing on the NBDC and JGA system enhancements.</p>]]></description>
            <pubDate><![CDATA[2020-11-06T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[MeDAS: a Metazoan Developmental Alternative Splicing database]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740539667-0e398349-7326-4859-b947-4fda87d0aa52/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa886</link>
            <description><![CDATA[<p class="para" id="N65541">Alternative splicing is widespread throughout eukaryotic genomes and greatly increases transcriptomic diversity. Many alternative isoforms have functional roles in developmental processes and are precisely temporally regulated. To facilitate the study of alternative splicing in a developmental context, we created MeDAS, a Metazoan Developmental Alternative Splicing database. MeDAS is an added-value resource that re-analyses publicly archived RNA-seq libraries to provide quantitative data on alternative splicing events as they vary across the time course of development. It has broad temporal and taxonomic scope and is intended to assist the user in identifying trends in alternative splicing throughout development. To create MeDAS, we re-analysed a curated set of 2232 Illumina polyA+ RNA-seq libraries that chart detailed time courses of embryonic and post-natal development across 18 species with a taxonomic range spanning the major metazoan lineages from <i>Caenorhabditis elegans</i> to human. MeDAS is freely available at https://das.chenlulab.com both as raw data tables and as an interactive browser allowing searches by species, tissue, or genomic feature (gene, transcript or exon ID and sequence). Results will provide details on alternative splicing events identified for the queried feature and can be visualised at the gene-, transcript- and exon-level as time courses of expression and inclusion levels, respectively.</p>]]></description>
            <pubDate><![CDATA[2020-10-21T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[VIPERdb v3.0: a structure-based data analytics platform for viral capsids]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740525192-2eddd228-fe46-45ac-a66f-d3910bdc3b0a/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1096</link>
            <description><![CDATA[<p class="para" id="N65541">VIrus Particle ExploreR data base (VIPERdb) (http://viperdb.scripps.edu) is a curated repository of virus capsid structures and a database of structure-derived data along with various virus specific information. VIPERdb has been continuously improved for over 20 years and contains a number of virus structure analysis tools. The release of VIPERdb v3.0 contains new structure-based data analytics tools like Multiple Structure-based and Sequence Alignment (MSSA) to identify hot-spot residues within a selected group of structures and an anomaly detection application to analyze and curate the structure-derived data within individual virus families. At the time of this writing, there are 931 virus structures from 62 different virus families in the database. Significantly, the new release also contains a standalone database called ‘Virus World database’ (VWdb) that comprises all the characterized viruses (∼181 000) known to date, gathered from ICTVdb and NCBI, and their capsid protein sequences, organized according to their virus taxonomy with links to known structures in VIPERdb and PDB. Moreover, the new release of VIPERdb includes a service-oriented data engine to handle all the data access requests and provides an interface for futuristic data analytics using machine leaning applications.</p>]]></description>
            <pubDate><![CDATA[2020-12-11T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Database Resources of the National Genomics Data Center, China National Center for Bioinformation in 2021]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740486800-a4fcd12d-3735-40c5-88af-4b626e4f9936/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1022</link>
            <description><![CDATA[<p class="para" id="N65541">The National Genomics Data Center (NGDC), part of the China National Center for Bioinformation (CNCB), provides a suite of database resources to support worldwide research activities in both academia and industry. With the explosive growth of multi-omics data, CNCB-NGDC is continually expanding, updating and enriching its core database resources through big data deposition, integration and translation. In the past year, considerable efforts have been devoted to 2019nCoVR, a newly established resource providing a global landscape of SARS-CoV-2 genomic sequences, variants, and haplotypes, as well as Aging Atlas, BrainBase, GTDB (Glycosyltransferases Database), LncExpDB, and TransCirc (Translation potential for circular RNAs). Meanwhile, a series of resources have been updated and improved, including BioProject, BioSample, GWH (Genome Warehouse), GVM (Genome Variation Map), GEN (Gene Expression Nebulas) as well as several biodiversity and plant resources. Particularly, BIG Search, a scalable, one-stop, cross-database search engine, has been significantly updated by providing easy access to a large number of internal and external biological resources from CNCB-NGDC, our partners, EBI and NCBI. All of these resources along with their services are publicly accessible at https://bigd.big.ac.cn.</p>]]></description>
            <pubDate><![CDATA[2020-11-11T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[DEG 15, an update of the Database of Essential Genes that includes built-in analysis tools]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740478659-2bd8716f-53bd-4580-936a-3b2606386e29/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa917</link>
            <description><![CDATA[<p class="para" id="N65541">Essential genes refer to genes that are required by an organism to survive under specific conditions. Studies of the minimal-gene-set for bacteria have elucidated fundamental cellular processes that sustain life. The past five years have seen a significant progress in identifying human essential genes, primarily due to the successful use of CRISPR/Cas9 in various types of human cells. DEG 15, a new release of the Database of Essential Genes (www.essentialgene.org), has provided major advancements, compared to DEG 10. Specifically, the number of eukaryotic essential genes has increased by more than fourfold, and that of prokaryotic ones has more than doubled. Of note, the human essential-gene number has increased by more than tenfold. Moreover, we have developed built-in analysis modules by which users can perform various analyses, such as essential-gene distributions between bacterial leading and lagging strands, sub-cellular localization distribution, enrichment analysis of gene ontology and KEGG pathways, and generation of Venn diagrams to compare and contrast gene sets between experiments. Additionally, the database offers customizable BLAST tools for performing species- and experiment-specific BLAST searches. Therefore, DEG comprehensively harbors updated human-curated essential-gene records among prokaryotes and eukaryotes with built-in tools to enhance essential-gene analysis.</p>]]></description>
            <pubDate><![CDATA[2020-10-23T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[iCSDB: an integrated database of CRISPR screens]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740441572-c8c5a9c4-6b9e-4bfa-be3f-d3b6b3b11af1/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa989</link>
            <description><![CDATA[<p class="para" id="N65541">High-throughput screening based on CRISPR-Cas9 libraries has become an attractive and powerful technique to identify target genes for functional studies. However, accessibility of public data is limited due to the lack of user-friendly utilities and up-to-date resources covering experiments from third parties. Here, we describe iCSDB, an integrated database of CRISPR screening experiments using human cell lines. We compiled two major sources of CRISPR-Cas9 screening: the DepMap portal and BioGRID ORCS. DepMap portal itself is an integrated database that includes three large-scale projects of CRISPR screening. We additionally aggregated CRISPR screens from BioGRID ORCS that is a collection of screening results from PubMed articles. Currently, iCSDB contains 1375 genome-wide screens across 976 human cell lines, covering 28 tissues and 70 cancer types. Importantly, the batch effects from different CRISPR libraries were removed and the screening scores were converted into a single metric to estimate the knockout efficiency. Clinical and molecular information were also integrated to help users to select cell lines of interest readily. Furthermore, we have implemented various interactive tools and viewers to facilitate users to choose, examine and compare the screen results both at the gene and guide RNA levels. iCSDB is available at https://www.kobic.re.kr/icsdb/.</p>]]></description>
            <pubDate><![CDATA[2020-11-02T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[FlyBase: updates to the <i>Drosophila melanogaster</i> knowledge base]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740397511-b4b0e084-e838-41a2-892b-e810783a8cb7/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1026</link>
            <description><![CDATA[<p class="para" id="N65541">FlyBase (flybase.org) is an essential online database for researchers using <i>Drosophila melanogaster</i> as a model organism, facilitating access to a diverse array of information that includes genetic, molecular, genomic and reagent resources. Here, we describe the introduction of several new features at FlyBase, including Pathway Reports, paralog information, disease models based on orthology, customizable tables within reports and overview displays (‘ribbons’) of expression and disease data. We also describe a variety of recent important updates, including incorporation of a developmental proteome, upgrades to the GAL4 search tab, additional Experimental Tool Reports, migration to JBrowse for genome browsing and improvements to batch queries/downloads and the Fast-Track Your Paper tool.</p>]]></description>
            <pubDate><![CDATA[2020-11-21T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[AcrHub: an integrative hub for investigating, predicting and mapping anti-CRISPR proteins]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740376409-71636662-9b8e-43b4-ba9c-62874ec6c071/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa951</link>
            <description><![CDATA[<p class="para" id="N65541">Anti-CRISPR (Acr) proteins naturally inhibit CRISPR-Cas adaptive immune systems across bacterial and archaeal domains of life. This emerging field has caused a paradigm shift in the way we think about the CRISPR-Cas system, and promises a number of useful applications from gene editing to phage therapy. As the number of verified and predicted Acrs rapidly expands, few online resources have been developed to deal with this wealth of information. To overcome this shortcoming, we developed AcrHub, an integrative database to provide an all-in-one solution for investigating, predicting and mapping Acr proteins. AcrHub catalogs 339 non-redundant experimentally validated Acrs and over 70 000 predicted Acrs extracted from genome sequence data from a diverse range of prokaryotic organisms and their viruses. It integrates state-of-the-art predictors to predict potential Acrs, and incorporates three analytical modules: similarity analysis, phylogenetic analysis and homology network analysis, to analyze their relationships with known Acrs. By interconnecting all modules as a platform, AcrHub presents enriched and in-depth analysis of known and potential Acrs and therefore provides new and exciting insights into the future of Acr discovery and validation. AcrHub is freely available at http://pacrispr.erc.monash.edu/AcrHub/.</p><p class="para" id="N65542">
<div class="section" id="ga1"><div class="img"><div class="imgeVideo"><div class="img-fullscreenIcon" onClick="javascript:showImageContent('ga1');"><img src="/public/images/journalImg/fullscreen.png"/></div><div class="imageVideo"><img src="/dataresources/secured/content-1765740376409-71636662-9b8e-43b4-ba9c-62874ec6c071/assets/gkaa951gra1.jpg" alt="AcrHub: an integrative hub for investigating, predicting and mapping anti-CRISPR proteins."/></div></div><div class="imgeVideoCaption" id="N65544"><div class="captionTitle">Graphical Abstract</div><div class="captionText">                                      AcrHub: an integrative hub for investigating, predicting and mapping anti-CRISPR proteins.</div></div></div></div>
</p>]]></description>
            <pubDate><![CDATA[2020-11-02T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[FANTOM enters 20th year: expansion of transcriptomic atlases and functional annotation of non-coding RNAs]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740362021-b5a53e86-954f-41ec-9a13-9943955ba769/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1054</link>
            <description><![CDATA[<p class="para" id="N65541">The Functional ANnoTation Of the Mammalian genome (FANTOM) Consortium has continued to provide extensive resources in the pursuit of understanding the transcriptome, and transcriptional regulation, of mammalian genomes for the last 20 years. To share these resources with the research community, the FANTOM web-interfaces and databases are being regularly updated, enhanced and expanded with new data types. In recent years, the FANTOM Consortium's efforts have been mainly focused on creating new non-coding RNA datasets and resources. The existing FANTOM5 human and mouse miRNA atlas was supplemented with rat, dog, and chicken datasets. The sixth (latest) edition of the FANTOM project was launched to assess the function of human long non-coding RNAs (lncRNAs). From its creation until 2020, FANTOM6 has contributed to the research community a large dataset generated from the knock-down of 285 lncRNAs in human dermal fibroblasts; this is followed with extensive expression profiling and cellular phenotyping. Other updates to the FANTOM resource includes the reprocessing of the miRNA and promoter atlases of human, mouse and chicken with the latest reference genome assemblies. To facilitate the use and accessibility of all above resources we further enhanced FANTOM data viewers and web interfaces. The updated FANTOM web resource is publicly available at https://fantom.gsc.riken.jp/.</p>]]></description>
            <pubDate><![CDATA[2020-11-19T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[WikiPathways: connecting communities]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740348624-e215afb1-7110-48f6-95d2-0202e8109daf/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1024</link>
            <description><![CDATA[<p class="para" id="N65541">WikiPathways (https://www.wikipathways.org) is a biological pathway database known for its collaborative nature and open science approaches. With the core idea of the scientific community developing and curating biological knowledge in pathway models, WikiPathways lowers all barriers for accessing and using its content. Increasingly more content creators, initiatives, projects and tools have started using WikiPathways. Central in this growth and increased use of WikiPathways are the various communities that focus on particular subsets of molecular pathways such as for rare diseases and lipid metabolism. Knowledge from published pathway figures helps prioritize pathway development, using optical character and named entity recognition. We show the growth of WikiPathways over the last three years, highlight the new communities and collaborations of pathway authors and curators, and describe various technologies to connect to external resources and initiatives. The road toward a sustainable, community-driven pathway database goes through integration with other resources such as Wikidata and allowing more use, curation and redistribution of WikiPathways content.</p><p class="para" id="N65542">
<div class="section" id="ga1"><div class="img"><div class="imgeVideo"><div class="img-fullscreenIcon" onClick="javascript:showImageContent('ga1');"><img src="/public/images/journalImg/fullscreen.png"/></div><div class="imageVideo"><img src="/dataresources/secured/content-1765740348624-e215afb1-7110-48f6-95d2-0202e8109daf/assets/gkaa1024gra1.jpg" alt="WikiPathways enables research communities to collaborate on molecular pathway curation to create reusable, machine-readable pathway models. The pathway collections are freely available and integrated with many analysis tools and resources."/></div></div><div class="imgeVideoCaption" id="N65544"><div class="captionTitle">Graphical Abstract</div><div class="captionText">                                      WikiPathways enables research communities to collaborate on molecular pathway curation to create reusable, machine-readable pathway models. The pathway collections are freely available and integrated with many analysis tools and resources.</div></div></div></div>
</p>]]></description>
            <pubDate><![CDATA[2020-11-19T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[m<sup>6</sup>A-Atlas: a comprehensive knowledgebase for unraveling the <i>N</i><sup>6</sup>-methyladenosine (m<sup>6</sup>A) epitranscriptome]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740342935-f20640df-9ea4-4dbb-bcde-f6dff61be481/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa692</link>
            <description><![CDATA[<p class="para" id="N65541">
<i>N</i>
<sup>6</sup>-Methyladenosine (m<sup>6</sup>A) is the most prevalent RNA modification on mRNAs and lncRNAs. It plays a pivotal role during various biological processes and disease pathogenesis. We present here a comprehensive knowledgebase, m<sup>6</sup>A-Atlas, for unraveling the m<sup>6</sup>A epitranscriptome. Compared to existing databases, m<sup>6</sup>A-Atlas features a high-confidence collection of 442 162 reliable m<sup>6</sup>A sites identified from seven base-resolution technologies and the quantitative (rather than binary) epitranscriptome profiles estimated from 1363 high-throughput sequencing samples. It also offers novel features, such as; the conservation of m<sup>6</sup>A sites among seven vertebrate species (including human, mouse and chimp), the m<sup>6</sup>A epitranscriptomes of 10 virus species (including HIV, KSHV and DENV), the putative biological functions of individual m<sup>6</sup>A sites predicted from epitranscriptome data, and the potential pathogenesis of m<sup>6</sup>A sites inferred from disease-associated genetic mutations that can directly destroy m<sup>6</sup>A directing sequence motifs. A user-friendly graphical user interface was constructed to support the query, visualization and sharing of the m<sup>6</sup>A epitranscriptomes annotated with sites specifying their interaction with post-transcriptional machinery (RBP-binding, microRNA interaction and splicing sites) and interactively display the landscape of multiple RNA modifications. These resources provide fresh opportunities for unraveling the m<sup>6</sup>A epitranscriptomes. m<sup>6</sup>A-Atlas is freely accessible at: www.xjtlu.edu.cn/biologicalsciences/atlas.</p>]]></description>
            <pubDate><![CDATA[2020-08-21T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[KEGG: integrating viruses and cellular organisms]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740250651-4813e5ff-6059-4a95-b89a-a5a5bd518d21/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa970</link>
            <description><![CDATA[<p class="para" id="N65541">KEGG (https://www.kegg.jp/) is a manually curated resource integrating eighteen databases categorized into systems, genomic, chemical and health information. It also provides KEGG mapping tools, which enable understanding of cellular and organism-level functions from genome sequences and other molecular datasets. KEGG mapping is a predictive method of reconstructing molecular network systems from molecular building blocks based on the concept of functional orthologs. Since the introduction of the KEGG NETWORK database, various diseases have been associated with network variants, which are perturbed molecular networks caused by human gene variants, viruses, other pathogens and environmental factors. The network variation maps are created as aligned sets of related networks showing, for example, how different viruses inhibit or activate specific cellular signaling pathways. The KEGG pathway maps are now integrated with network variation maps in the NETWORK database, as well as with conserved functional units of KEGG modules and reaction modules in the MODULE database. The KO database for functional orthologs continues to be improved and virus KOs are being expanded for better understanding of virus-cell interactions and for enabling prediction of viral perturbations.</p>]]></description>
            <pubDate><![CDATA[2020-10-30T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Lnc2Cancer 3.0: an updated resource for experimentally supported lncRNA/circRNA cancer associations and web tools based on RNA-seq and scRNA-seq data]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740191574-be8a3a74-08a1-4db1-a410-5a8ff3b75b3c/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1006</link>
            <description><![CDATA[<p class="para" id="N65541">An updated Lnc2Cancer 3.0 (http://www.bio-bigdata.net/lnc2cancer or http://bio-bigdata.hrbmu.edu.cn/lnc2cancer) database, which includes comprehensive data on experimentally supported long non-coding RNAs (lncRNAs) and circular RNAs (circRNAs) associated with human cancers. In addition, web tools for analyzing lncRNA expression by high-throughput RNA sequencing (RNA-seq) and single-cell RNA-seq (scRNA-seq) are described. Lnc2Cancer 3.0 was updated with several new features, including (i) Increased cancer-associated lncRNA entries over the previous version. The current release includes 9254 lncRNA-cancer associations, with 2659 lncRNAs and 216 cancer subtypes. (ii) Newly adding 1049 experimentally supported circRNA-cancer associations, with 743 circRNAs and 70 cancer subtypes. (iii) Experimentally supported regulatory mechanisms of cancer-related lncRNAs and circRNAs, involving microRNAs, transcription factors (TF), genetic variants, methylation and enhancers were included. (iv) Appending experimentally supported biological functions of cancer-related lncRNAs and circRNAs including cell growth, apoptosis, autophagy, epithelial mesenchymal transformation (EMT), immunity and coding ability. (v) Experimentally supported clinical relevance of cancer-related lncRNAs and circRNAs in metastasis, recurrence, circulation, drug resistance, and prognosis was included. Additionally, two flexible online tools, including RNA-seq and scRNA-seq web tools, were developed to enable fast and customizable analysis and visualization of lncRNAs in cancers. Lnc2Cancer 3.0 is a valuable resource for elucidating the associations between lncRNA, circRNA and cancer.</p>]]></description>
            <pubDate><![CDATA[2020-11-21T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[HeRA: an atlas of enhancer RNAs across human tissues]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740159797-c79dbdc4-4d18-4589-b42d-ade8134fb633/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa940</link>
            <description><![CDATA[<p class="para" id="N65541">Enhancer RNA (eRNA) is a type of long non-coding RNA transcribed from DNA enhancer regions. Despite critical roles of eRNA in gene regulation, the expression landscape of eRNAs in normal human tissue remains unexplored. Using numerous samples from the Genotype-Tissue Expression project, we characterized 45 411 detectable eRNAs and identified tens of thousands of associations between eRNAs and traits, including gender, race, and age. We constructed a co-expression network to identify millions of putative eRNA regulators and target genes across different tissues. We further constructed a user-friendly data portal, <span style="text-decoration: underline">H</span>uman <span style="text-decoration: underline">e</span>nhancer <span style="text-decoration: underline">R</span>NA <span style="text-decoration: underline">A</span>tlas (HeRA, https://hanlab.uth.edu/HeRA/). In HeRA, users can search, browse, and download the eRNA expression profile, trait-related eRNAs, and eRNA co-expression network by searching the eRNA ID, gene symbol, and genomic region in one or multiple tissues. HeRA is the first data portal to characterize eRNAs from 9577 samples across 54 human tissues and facilitates functional and mechanistic investigations of eRNAs.</p>]]></description>
            <pubDate><![CDATA[2020-10-29T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[KinaseMD: kinase mutations and drug response database]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740132726-4d4d167c-5bad-4821-9258-a575c283082d/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa945</link>
            <description><![CDATA[<p class="para" id="N65541">Mutations in kinases are abundant and critical to study signaling pathways and regulatory roles in human disease, especially in cancer. Somatic mutations in kinase genes can affect drug treatment, both sensitivity and resistance, to clinically used kinase inhibitors. Here, we present a newly constructed database, KinaseMD (<b>kinase m</b>utations and <b>d</b>rug response), to structurally and functionally annotate kinase mutations. KinaseMD integrates 679 374 somatic mutations, 251 522 network-rewiring events, and 390 460 drug response records curated from various sources for 547 kinases. We uniquely annotate the mutations and kinase inhibitor response in four types of protein substructures (gatekeeper, A-loop, G-loop and αC-helix) that are linked to kinase inhibitor resistance in literature. In addition, we annotate functional mutations that may rewire kinase regulatory network and report four phosphorylation signals (gain, loss, up-regulation and down-regulation). Overall, KinaseMD provides the most updated information on mutations, unique annotations of drug response especially drug resistance and functional sites of kinases. KinaseMD is accessible at https://bioinfo.uth.edu/kmd/, having functions for searching, browsing and downloading data. To our knowledge, there has been no systematic annotation of these structural mutations linking to kinase inhibitor response. In summary, KinaseMD is a centralized database for kinase mutations and drug response.</p>]]></description>
            <pubDate><![CDATA[2020-11-02T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Datanator: an integrated database of molecular data for quantitatively modeling cellular behavior]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740075128-44c26fea-132f-4d93-8918-6eaf3abcc8f0/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1008</link>
            <description><![CDATA[<p class="para" id="N65541">Integrative research about multiple biochemical subsystems has significant potential to help advance biology, bioengineering and medicine. However, it is difficult to obtain the diverse data needed for integrative research. To facilitate biochemical research, we developed Datanator (https://datanator.info), an integrated database and set of tools for finding <i>clouds</i> of multiple types of molecular data about specific molecules and reactions in specific organisms and environments, as well as data about chemically-similar molecules and reactions in phylogenetically-similar organisms in similar environments. Currently, Datanator includes metabolite concentrations, RNA modifications and half-lives, protein abundances and modifications, and reaction rate constants about a broad range of organisms. Going forward, we aim to launch a community initiative to curate additional data. Datanator also provides tools for filtering, visualizing and exporting these data clouds. We believe that Datanator can facilitate a wide range of research from integrative mechanistic models, such as whole-cell models, to comparative data-driven analyses of multiple organisms.</p><p class="para" id="N65542">
<div class="section" id="ga1"><div class="img"><div class="imgeVideo"><div class="img-fullscreenIcon" onClick="javascript:showImageContent('ga1');"><img src="/public/images/journalImg/fullscreen.png"/></div><div class="imageVideo"><img src="/dataresources/secured/content-1765740075128-44c26fea-132f-4d93-8918-6eaf3abcc8f0/assets/gkaa1008gra1.jpg" alt="Datanator (https://datanator.info) is an integrated database of molecular data. To help investigators find relevant data for research even in the absence of direct measurements, Datanator includes tools for assembling &quot;clouds&quot; of measurements centered on a specific molecule or reaction in a particular organism that encompass measurements of similar molecules and reactions in similar organisms. Datanator is ideal for integrative and comparative analysis and modeling."/></div></div><div class="imgeVideoCaption" id="N65544"><div class="captionTitle">Graphical Abstract</div><div class="captionText">                                      Datanator (https://datanator.info) is an integrated database of molecular data. To help investigators find relevant data for research even in the absence of direct measurements, Datanator includes tools for assembling "clouds" of measurements centered on a specific molecule or reaction in a particular organism that encompass measurements of similar molecules and reactions in similar organisms. Datanator is ideal for integrative and comparative analysis and modeling.</div></div></div></div>
</p>]]></description>
            <pubDate><![CDATA[2020-11-11T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[MobiDB: intrinsically disordered proteins in 2021]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740015420-47cf02b4-c311-4e01-9e8a-10ad91bfb571/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1058</link>
            <description><![CDATA[<p class="para" id="N65541">The MobiDB database (URL: https://mobidb.org/) provides predictions and annotations for intrinsically disordered proteins. Here, we report recent developments implemented in MobiDB version 4, regarding the database format, with novel types of annotations and an improved update process. The new website includes a re-designed user interface, a more effective search engine and advanced API for programmatic access. The new database schema gives more flexibility for the users, as well as simplifying the maintenance and updates. In addition, the new entry page provides more visualisation tools including customizable feature viewer and graphs of the residue contact maps. MobiDB v4 annotates the binding modes of disordered proteins, whether they undergo disorder-to-order transitions or remain disordered in the bound state. In addition, disordered regions undergoing liquid-liquid phase separation or post-translational modifications are defined. The integrated information is presented in a simplified interface, which enables faster searches and allows large customized datasets to be downloaded in TSV, Fasta or JSON formats. An alternative advanced interface allows users to drill deeper into features of interest. A new statistics page provides information at database and proteome levels. The new MobiDB version presents state-of-the-art knowledge on disordered proteins and improves data accessibility for both computational and experimental users.</p>]]></description>
            <pubDate><![CDATA[2020-11-25T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Animal-APAdb: a comprehensive animal alternative polyadenylation database]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739980054-4b88d68d-1383-40eb-a4c7-1d2e57fe5c16/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa778</link>
            <description><![CDATA[<p class="para" id="N65541">Alternative polyadenylation (APA) is an important post-transcriptional regulatory mechanism that recognizes different polyadenylation signals on transcripts, resulting in transcripts with different lengths of 3′ untranslated regions and thereby influencing a series of biological processes. Recent studies have highlighted the important roles of APA in human. However, APA profiles in other animals have not been fully recognized, and there is no database that provides comprehensive APA information for other animals except human. Here, by using the RNA sequencing data collected from public databases, we systematically characterized the APA profiles in 9244 samples of 18 species. In total, we identified 342 952 APA events with a median of 17 020 per species using the DaPars2 algorithm, and 315 691 APA events with a median of 17 953 per species using the QAPA algorithm in these 18 species, respectively. In addition, we predicted the polyadenylation sites (PAS) and motifs near PAS of these species. We further developed Animal-APAdb, a user-friendly database (http://gong_lab.hzau.edu.cn/Animal-APAdb/) for data searching, browsing and downloading. With comprehensive information of APA events in different tissues of different species, Animal-APAdb may greatly facilitate the exploration of animal APA patterns and novel mechanisms, gene expression regulation and APA evolution across tissues and species.</p>]]></description>
            <pubDate><![CDATA[2020-09-28T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The UCSC Genome Browser database: 2021 update]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739973171-a6481bd2-41f3-4115-9ae9-f962ffe7bd8c/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1070</link>
            <description><![CDATA[<p class="para" id="N65541">For more than two decades, the UCSC Genome Browser database (https://genome.ucsc.edu) has provided high-quality genomics data visualization and genome annotations to the research community. As the field of genomics grows and more data become available, new modes of display are required to accommodate new technologies. New features released this past year include a Hi-C heatmap display, a phased family trio display for VCF files, and various track visualization improvements. Striving to keep data up-to-date, new updates to gene annotations include GENCODE Genes, NCBI RefSeq Genes, and Ensembl Genes. New data tracks added for human and mouse genomes include the ENCODE registry of candidate cis-regulatory elements, promoters from the Eukaryotic Promoter Database, and NCBI RefSeq Select and Matched Annotation from NCBI and EMBL-EBI (MANE). Within weeks of learning about the outbreak of coronavirus, UCSC released a genome browser, with detailed annotation tracks, for the SARS-CoV-2 RNA reference assembly.</p>]]></description>
            <pubDate><![CDATA[2020-11-22T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[PROMISCUOUS 2.0: a resource for drug-repositioning]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739930034-a11fcbef-98e7-464c-8f22-07965ccb897e/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1061</link>
            <description><![CDATA[<p class="para" id="N65541">The development of new drugs for diseases is a time-consuming, costly and risky process. In recent years, many drugs could be approved for other indications. This repurposing process allows to effectively reduce development costs, time and, ultimately, save patients’ lives. During the ongoing COVID-19 pandemic, drug repositioning has gained widespread attention as a fast opportunity to find potential treatments against the newly emerging disease. In order to expand this field to researchers with varying levels of experience, we made an effort to open it to all users (meaning novices as well as experts in cheminformatics) by significantly improving the entry-level user experience. The browsing functionality can be used as a global entry point to collect further information with regards to small molecules (∼1 million), side-effects (∼110 000) or drug-target interactions (∼3 million). The drug-repositioning tab for small molecules will also suggest possible drug-repositioning opportunities to the user by using structural similarity measurements for small molecules using two different approaches. Additionally, using information from the Promiscuous 2.0 Database, lists of candidate drugs for given indications were precomputed, including a section dedicated to potential treatments for COVID-19. All the information is interconnected by a dynamic network-based visualization to identify new indications for available compounds. Promiscuous 2.0 is unique in its functionality and is publicly available at http://bioinformatics.charite.de/promiscuous2.</p>]]></description>
            <pubDate><![CDATA[2020-11-16T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Aging Atlas: a multi-omics database for aging biology]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739906612-8d707836-fbd8-4a1f-89af-8a6f3f7a1aa4/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa894</link>
            <description><![CDATA[<p class="para" id="N65541">Organismal aging is driven by interconnected molecular changes encompassing internal and extracellular factors. Combinational analysis of high-throughput ‘multi-omics’ datasets (gathering information from genomics, epigenomics, transcriptomics, proteomics, metabolomics and pharmacogenomics), at either populational or single-cell levels, can provide a multi-dimensional, integrated profile of the heterogeneous aging process with unprecedented throughput and detail. These new strategies allow for the exploration of the molecular profile and regulatory status of gene expression during aging, and in turn, facilitate the development of new aging interventions. With a continually growing volume of valuable aging-related data, it is necessary to establish an open and integrated database to support a wide spectrum of aging research. The Aging Atlas database aims to provide a wide range of life science researchers with valuable resources that allow access to a large-scale of gene expression and regulation datasets created by various high-throughput omics technologies. The current implementation includes five modules: transcriptomics (RNA-seq), single-cell transcriptomics (scRNA-seq), epigenomics (ChIP-seq), proteomics (protein–protein interaction), and pharmacogenomics (geroprotective compounds). Aging Atlas provides user-friendly functionalities to explore age-related changes in gene expression, as well as raw data download services. Aging Atlas is freely available at https://bigd.big.ac.cn/aging/index.</p>]]></description>
            <pubDate><![CDATA[2020-10-29T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[RMVar: an updated database of functional variants involved in RNA modifications]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739901852-0025c90e-3c66-4c8e-b094-bbfe71352c8f/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa811</link>
            <description><![CDATA[<p class="para" id="N65541">Distinguishing the few disease-related variants from a massive number of passenger variants is a major challenge. Variants affecting RNA modifications that play critical roles in many aspects of RNA metabolism have recently been linked to many human diseases, such as cancers. Evaluating the effect of genetic variants on RNA modifications will provide a new perspective for understanding the pathogenic mechanism of human diseases. Previously, we developed a database called ‘m6AVar’ to host variants associated with m<sup>6</sup>A, one of the most prevalent RNA modifications in eukaryotes. To host all RNA modification (RM)-associated variants, here we present an updated version of m6AVar renamed RMVar (http://rmvar.renlab.org). In this update, RMVar contains 1 678 126 RM-associated variants for 9 kinds of RNA modifications, namely m<sup>6</sup>A, m<sup>6</sup>Am, m<sup>1</sup>A, pseudouridine, m<sup>5</sup>C, m<sup>5</sup>U, 2′-O-Me, A-to-I and m<sup>7</sup>G, at three confidence levels. Moreover, RBP binding regions, miRNA targets, splicing events and circRNAs were integrated to assist investigations of the effects of RM-associated variants on posttranscriptional regulation. In addition, disease-related information was integrated from ClinVar and other genome-wide association studies (GWAS) to investigate the relationship between RM-associated variants and diseases. We expect that RMVar may boost further functional studies on genetic variants affecting RNA modifications.</p><p class="para" id="N65542">
<div class="section" id="ga1"><div class="img"><div class="imgeVideo"><div class="img-fullscreenIcon" onClick="javascript:showImageContent('ga1');"><img src="/public/images/journalImg/fullscreen.png"/></div><div class="imageVideo"><img src="/dataresources/secured/content-1765739901852-0025c90e-3c66-4c8e-b094-bbfe71352c8f/assets/gkaa811gra1.jpg" alt="RMVar: an updated database of functional variants involved in RNA modifications."/></div></div><div class="imgeVideoCaption" id="N65544"><div class="captionTitle">Graphical Abstract</div><div class="captionText">                                      RMVar: an updated database of functional variants involved in RNA modifications.</div></div></div></div>
</p>]]></description>
            <pubDate><![CDATA[2020-10-06T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[IDDB: a comprehensive resource featuring genes, variants and characteristics associated with infertility]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739897338-188483d2-443e-4bd8-b5e6-2bad1606844e/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa753</link>
            <description><![CDATA[<p class="para" id="N65541">Infertility is a complex multifactorial disease that affects up to 10% of couples across the world. However, many mechanisms of infertility remain unclear due to the lack of studies based on systematic knowledge, leading to ineffective treatment and/or transmission of genetic defects to offspring. Here, we developed an infertility disease database to provide a comprehensive resource featuring various factors involved in infertility. Features in the current IDDB version were manually curated as follows: (i) a total of 307 infertility-associated genes in human and 1348 genes associated with reproductive disorder in 9 model organisms; (ii) a total of 202 chromosomal abnormalities leading to human infertility, including aneuploidies and structural variants; and (iii) a total of 2078 pathogenic variants from infertility patients’ samples across 60 different diseases causing infertility. Additionally, the characteristics of clinically diagnosed infertility patients (i.e. causative variants, laboratory indexes and clinical manifestations) were collected. To the best of our knowledge, the IDDB is the first infertility database serving as a systematic resource for biologists to decipher infertility mechanisms and for clinicians to achieve better diagnosis/treatment of patients from disease phenotype to genetic factors. The IDDB is freely available at http://mdl.shsmu.edu.cn/IDDB/.</p><p class="para" id="N65542">
<div class="section" id="ga1"><div class="img"><div class="imgeVideo"><div class="img-fullscreenIcon" onClick="javascript:showImageContent('ga1');"><img src="/public/images/journalImg/fullscreen.png"/></div><div class="imageVideo"><img src="/dataresources/secured/content-1765739897338-188483d2-443e-4bd8-b5e6-2bad1606844e/assets/gkaa753gra1.jpg" alt="The major components of Infertility Disease Database and their combinations."/></div></div><div class="imgeVideoCaption" id="N65544"><div class="captionTitle">Graphical Abstract</div><div class="captionText">                                      The major components of Infertility Disease Database and their combinations.</div></div></div></div>
</p>]]></description>
            <pubDate><![CDATA[2020-09-17T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[dbGuide: a database of functionally validated guide RNAs for genome editing in human and mouse cells]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739886291-acf794f4-6751-4427-affe-706c203b3e5d/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa848</link>
            <description><![CDATA[<p class="para" id="N65541">With the technology's accessibility and ease of use, CRISPR has been employed widely in many different organisms and experimental settings. As a result, thousands of publications have used CRISPR to make specific genetic perturbations, establishing in itself a resource of validated guide RNA sequences. While numerous computational tools to assist in the design and identification of candidate guide RNAs exist, these are still just at best predictions and generally, researchers inevitably will test multiple sequences for functional activity. Here, we present <i>dbGuide</i> (https://sgrnascorer.cancer.gov/dbguide), a database of functionally validated guide RNA sequences for CRISPR/Cas9-based knockout in human and mouse. Our database not only contains computationally determined candidate guide RNA sequences, but of even greater value, over 4000 sequences which have been functionally validated either through direct amplicon sequencing or manual curation of literature from over 1000 publications. Finally, our established framework will allow for continual addition of newly published and experimentally validated guide RNA sequences for CRISPR/Cas9-based knockout as well as incorporation of sequences from different gene editing systems, additional species and other types of site-specific functionalities such as base editing, gene activation, repression and epigenetic modification.</p>]]></description>
            <pubDate><![CDATA[2020-10-13T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[GRNdb: decoding the gene regulatory networks in diverse human and mouse conditions]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739822202-dae95753-a6e8-4d0f-adf8-f342de1b6047/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa995</link>
            <description><![CDATA[<p class="para" id="N65541">Gene regulatory networks (GRNs) formed by transcription factors (TFs) and their downstream target genes play essential roles in gene expression regulation. Moreover, GRNs can be dynamic changing across different conditions, which are crucial for understanding the underlying mechanisms of disease pathogenesis. However, no existing database provides comprehensive GRN information for various human and mouse normal tissues and diseases at the single-cell level. Based on the known TF-target relationships and the large-scale single-cell RNA-seq data collected from public databases as well as the bulk data of The Cancer Genome Atlas and the Genotype-Tissue Expression project, we systematically predicted the GRNs of 184 different physiological and pathological conditions of human and mouse involving &gt;633 000 cells and &gt;27 700 bulk samples. We further developed GRNdb, a freely accessible and user-friendly database (http://www.grndb.com/) for searching, comparing, browsing, visualizing, and downloading the predicted information of 77 746 GRNs, 19 687 841 TF-target pairs, and related binding motifs at single-cell/bulk resolution. GRNdb also allows users to explore the gene expression profile, correlations, and the associations between expression levels and the patient survival of diverse cancers. Overall, GRNdb provides a valuable and timely resource to the scientific community to elucidate the functions and mechanisms of gene expression regulation in various conditions.</p>]]></description>
            <pubDate><![CDATA[2020-11-05T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Plant-ImputeDB: an integrated multiple plant reference panel database for genotype imputation]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739790253-11f7d0a7-2e2f-45cc-9b35-4a054a5b0453/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa953</link>
            <description><![CDATA[<p class="para" id="N65541">Genotype imputation is a process that estimates missing genotypes in terms of the haplotypes and genotypes in a reference panel. It can effectively increase the density of single nucleotide polymorphisms (SNPs), boost the power to identify genetic association and promote the combination of genetic studies. However, there has been a lack of high-quality reference panels for most plants, which greatly hinders the application of genotype imputation. Here, we developed Plant-ImputeDB (http://gong_lab.hzau.edu.cn/Plant_imputeDB/), a comprehensive database with reference panels of 12 plant species for online genotype imputation, SNP and block search and free download. By integrating genotype data and whole-genome resequencing data of plants from various studies and databases, the current Plant-ImputeDB provides high-quality reference panels of 12 plant species, including ∼69.9 million SNPs from 34 244 samples. It also provides an easy-to-use online tool with the option of two popular tools specifically designed for genotype imputation. In addition, Plant-ImputeDB accepts submissions of different types of genomic variations, and provides free and open access to all publicly available data in support of related research worldwide. In general, Plant-ImputeDB may serve as an important resource for plant genotype imputation and greatly facilitate the research on plant genetic research.</p>]]></description>
            <pubDate><![CDATA[2020-11-02T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[INTEDE: interactome of drug-metabolizing enzymes]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739685351-047936ff-9375-47ec-83cf-fde4160c731c/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa755</link>
            <description><![CDATA[<p class="para" id="N65541">Drug-metabolizing enzymes (DMEs) are critical determinant of drug safety and efficacy, and the interactome of DMEs has attracted extensive attention. There are 3 major interaction types in an interactome: microbiome–DME interaction (MICBIO), xenobiotics–DME interaction (XEOTIC) and host protein–DME interaction (HOSPPI). The interaction data of each type are essential for drug metabolism, and the collective consideration of multiple types has implication for the future practice of precision medicine. However, no database was designed to systematically provide the data of all types of DME interactions. Here, a database of the Interactome of Drug-Metabolizing Enzymes (INTEDE) was therefore constructed to offer these interaction data. First, 1047 unique DMEs (448 host and 599 microbial) were confirmed, for the first time, using their metabolizing drugs. Second, for these newly confirmed DMEs, all types of their interactions (3359 MICBIOs between 225 microbial species and 185 DMEs; 47 778 XEOTICs between 4150 xenobiotics and 501 DMEs; 7849 HOSPPIs between 565 human proteins and 566 DMEs) were comprehensively collected and then provided, which enabled the crosstalk analysis among multiple types. Because of the huge amount of accumulated data, the INTEDE made it possible to generalize key features for revealing disease etiology and optimizing clinical treatment. INTEDE is freely accessible at: https://idrblab.org/intede/</p><p class="para" id="N65542">
<div class="section" id="ga1"><div class="img"><div class="imgeVideo"><div class="img-fullscreenIcon" onClick="javascript:showImageContent('ga1');"><img src="/public/images/journalImg/fullscreen.png"/></div><div class="imageVideo"><img src="/dataresources/secured/content-1765739685351-047936ff-9375-47ec-83cf-fde4160c731c/assets/gkaa755gra1.jpg" alt="The Interactome Database of Drug-Metabolizing Enzyme (INTEDE) provides the comprehensive interactome data from three perspectives for both human and microbial drug-metabolizing enzymes."/></div></div><div class="imgeVideoCaption" id="N65544"><div class="captionTitle">Graphical Abstract</div><div class="captionText">                                      The Interactome Database of Drug-Metabolizing Enzyme (INTEDE) provides the comprehensive interactome data from three perspectives for both human and microbial drug-metabolizing enzymes.</div></div></div></div>
</p>]]></description>
            <pubDate><![CDATA[2020-10-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[RJunBase: a database of RNA splice junctions in human normal and cancerous tissues]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739616010-e1313cc8-676b-4c44-b6ae-923b63924c8f/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1056</link>
            <description><![CDATA[<p class="para" id="N65541">Splicing is an essential step of RNA processing for multi-exon genes, in which introns are removed from a precursor RNA, thereby producing mature RNAs containing splice junctions. Here, we develope the RJunBase (www.RJunBase.org), a web-accessible database of three types of RNA splice junctions (linear, back-splice, and fusion junctions) that are derived from RNA-seq data of non-cancerous and cancerous tissues. The RJunBase aims to integrate and characterize all RNA splice junctions of both healthy or pathological human cells and tissues. This new database facilitates the visualization of the gene-level splicing pattern and the junction-level expression profile, as well as the demonstration of unannotated and tumor-specific junctions. The first release of RJunBase contains 682 017 linear junctions, 225 949 back-splice junctions and 34 733 fusion junctions across 18 084 non-cancerous and 11 540 cancerous samples. RJunBase can aid researchers in discovering new splicing-associated targets and provide insights into the identification and assessment of potential neoepitopes for cancer treatment.</p>]]></description>
            <pubDate><![CDATA[2020-11-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[OGEE v3: Online GEne Essentiality database with increased coverage of organisms and human cell lines]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739604025-1c879f6d-1ad5-45d8-ae3c-d40039ebf05b/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa884</link>
            <description><![CDATA[<p class="para" id="N65541">OGEE is an Online GEne Essentiality database. Gene essentiality is not a static and binary property, rather a context-dependent and evolvable property in all forms of life. In OGEE we collect not only experimentally tested essential and non-essential genes, but also associated gene properties that contributes to gene essentiality. We tagged conditionally essential genes that show variable essentiality statuses across datasets to highlight complex interplays between gene functions and environmental/experimental perturbations. OGEE v3 contains gene essentiality datasets for 91 species; almost doubled from 48 species in previous version. To accommodate recent advances on human cancer essential genes (as known as tumor dependency genes) that could serve as targets for cancer treatment and/or drug development, we expanded the collection of human essential genes from 16 cell lines in previous to 581. These human cancer cell lines were tested with high-throughput experiments such as CRISPR-Cas9 and RNAi; in total, 150 of which were tested by both techniques. We also included factors known to contribute to gene essentiality for these cell lines, such as genomic mutation, methylation and gene expression, along with extensive graphical visualizations for ease of understanding of these factors. OGEE v3 can be accessible freely at https://v3.ogee.info.</p>]]></description>
            <pubDate><![CDATA[2020-10-21T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[PK-DB: pharmacokinetics database for individualized and stratified computational modeling]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739568853-6d51a36f-37e5-4943-9103-4d2aa100a738/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa990</link>
            <description><![CDATA[<p class="para" id="N65541">A multitude of pharmacokinetics studies have been published. However, due to the lack of an open database, pharmacokinetics data, as well as the corresponding meta-information, have been difficult to access. We present PK-DB (https://pk-db.com), an open database for pharmacokinetics information from clinical trials. PK-DB provides curated information on (i) characteristics of studied patient cohorts and subjects (e.g. age, bodyweight, smoking status, genetic variants); (ii) applied interventions (e.g. dosing, substance, route of application); (iii) pharmacokinetic parameters (e.g. clearance, half-life, area under the curve) and (iv) measured pharmacokinetic time-courses. Key features are the representation of experimental errors, the normalization of measurement units, annotation of information to biological ontologies, calculation of pharmacokinetic parameters from concentration-time profiles, a workflow for collaborative data curation, strong validation rules on the data, computational access via a REST API as well as human access via a web interface. PK-DB enables meta-analysis based on data from multiple studies and data integration with computational models. A special focus lies on meta-data relevant for individualized and stratified computational modeling with methods like physiologically based pharmacokinetic (PBPK), pharmacokinetic/pharmacodynamic (PK/PD), or population pharmacokinetic (pop PK) modeling.</p>]]></description>
            <pubDate><![CDATA[2020-11-05T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[MolluscDB: an integrated functional and evolutionary genomics database for the hyper-diverse animal phylum Mollusca]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739550298-b4e187ca-4f31-4633-83dd-5615b1ba7bba/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa918</link>
            <description><![CDATA[<p class="para" id="N65541">Mollusca represents the second largest animal phylum but remains poorly explored from a genomic perspective. While the recent increase in genomic resources holds great promise for a deep understanding of molluscan biology and evolution, access and utilization of these resources still pose a challenge. Here, we present the first comprehensive molluscan genomics database, MolluscDB (http://mgbase.qnlm.ac), which compiles and integrates current molluscan genomic/transcriptomic resources and provides convenient tools for multi-level integrative and comparative genomic analyses. MolluscDB enables a systematic view of genomic information from various aspects, such as genome assembly statistics, genome phylogenies, fossil records, gene information, expression profiles, gene families, transcription factors, transposable elements and mitogenome organization information. Moreover, MolluscDB offers valuable customized datasets or resources, such as gene coexpression networks across various developmental stages and adult tissues/organs, core gene repertoires inferred for major molluscan lineages, and macrosynteny analysis for chromosomal evolution. MolluscDB presents an integrative and comprehensive genomics platform that will allow the molluscan community to cope with ever-growing genomic resources and will expedite new scientific discoveries for understanding molluscan biology and evolution.</p>]]></description>
            <pubDate><![CDATA[2020-10-22T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[GIMICA: host genetic and immune factors shaping human microbiota]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739544736-3a948232-4501-4bf5-ae9e-c7514b052c24/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa851</link>
            <description><![CDATA[<p class="para" id="N65541">Besides the environmental factors having tremendous impacts on the composition of microbial community, the host factors have recently gained extensive attentions on their roles in shaping human microbiota. There are two major types of host factors: host genetic factors (HGFs) and host immune factors (HIFs). These factors of each type are essential for defining the chemical and physical landscapes inhabited by microbiota, and the collective consideration of both types have great implication to serve comprehensive health management. However, no database was available to provide the comprehensive factors of both types. Herein, a database entitled ‘Host <span style="text-decoration: underline">G</span>enetic and <span style="text-decoration: underline">I</span>mmune Factors Shaping Human <span style="text-decoration: underline">Mic</span>robiot<span style="text-decoration: underline">a</span> (GIMICA)’ was constructed. Based on the 4257 microbes confirmed to inhabit nine sites of human body, 2851 HGFs (1368 single nucleotide polymorphisms (SNPs), 186 copy number variations (CNVs), and 1297 non-coding ribonucleic acids (RNAs)) modulating the expression of 370 microbes were collected, and 549 HIFs (126 lymphocytes and phagocytes, 387 immune proteins, and 36 immune pathways) regulating the abundance of 455 microbes were also provided. All in all, GIMICA enabled the collective consideration not only between different types of host factor but also between the host and environmental ones, which is freely accessible without login requirement at: https://idrblab.org/gimica/</p><p class="para" id="N65542">
<div class="section" id="ga1"><div class="img"><div class="imgeVideo"><div class="img-fullscreenIcon" onClick="javascript:showImageContent('ga1');"><img src="/public/images/journalImg/fullscreen.png"/></div><div class="imageVideo"><img src="/dataresources/secured/content-1765739544736-3a948232-4501-4bf5-ae9e-c7514b052c24/assets/gkaa851gra1.jpg" alt="GIMICA is unique in: (i) providing both host genetic and immune factors shaping human microbiota and (ii) enabling collective consideration not only among various host factors but also between the host and environmental ones."/></div></div><div class="imgeVideoCaption" id="N65544"><div class="captionTitle">Graphical Abstract</div><div class="captionText">                                      GIMICA is unique in: (i) providing both host genetic and immune factors shaping human microbiota and (ii) enabling collective consideration not only among various host factors but also between the host and environmental ones.</div></div></div></div>
</p>]]></description>
            <pubDate><![CDATA[2020-10-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[NONCODEV6: an updated database dedicated to long non-coding RNA annotation in both animals and plants]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739485091-b5ed246f-900c-4a1c-9204-a839d52157d0/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1046</link>
            <description><![CDATA[<p class="para" id="N65541">NONCODE (http://www.noncode.org/) is a comprehensive database of collection and annotation of noncoding RNAs, especially long non-coding RNAs (lncRNAs) in animals. NONCODEV6 is dedicated to providing the full scope of lncRNAs across plants and animals. The number of lncRNAs in NONCODEV6 has increased from 548 640 to 644 510 since the last update in 2017. The number of human lncRNAs has increased from 172 216 to 173 112. The number of mouse lncRNAs increased from 131 697 to 131 974. The number of plant lncRNAs is 94 697. The relationship between lncRNAs in human and cancer were updated with transcriptome sequencing profiles. Three important new features were also introduced in NONCODEV6: (i) updated human lncRNA-disease relationships, especially cancer; (ii) lncRNA annotations with tissue expression profiles and predicted function in five common plants; iii) lncRNAs conservation annotation at transcript level for 23 plant species. NONCODEV6 is accessible through http://www.noncode.org/.</p>]]></description>
            <pubDate><![CDATA[2020-11-16T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[HERB: a high-throughput experiment- and reference-guided database of traditional Chinese medicine]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739431085-218c363e-0032-496e-be2e-c1d77e68d77b/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1063</link>
            <description><![CDATA[<p class="para" id="N65541">Pharmacotranscriptomics has become a powerful approach for evaluating the therapeutic efficacy of drugs and discovering new drug targets. Recently, studies of traditional Chinese medicine (TCM) have increasingly turned to high-throughput transcriptomic screens for molecular effects of herbs/ingredients. And numerous studies have examined gene targets for herbs/ingredients, and link herbs/ingredients to various modern diseases. However, there is currently no systematic database organizing these data for TCM. Therefore, we built <b>HERB</b>, a <b><span style="text-decoration: underline">h</span></b>igh-throughput <b><span style="text-decoration: underline">e</span></b>xperiment- and <b><span style="text-decoration: underline">r</span></b>eference-guided data<b><span style="text-decoration: underline">b</span></b>ase of TCM, with its Chinese name as BenCaoZuJian. We re-analyzed 6164 gene expression profiles from 1037 high-throughput experiments evaluating TCM herbs/ingredients, and generated connections between TCM herbs/ingredients and 2837 modern drugs by mapping the comprehensive pharmacotranscriptomics dataset in HERB to CMap, the largest such dataset for modern drugs. Moreover, we manually curated 1241 gene targets and 494 modern diseases for 473 herbs/ingredients from 1966 references published recently, and cross-referenced this novel information to databases containing such data for drugs. Together with database mining and statistical inference, we linked 12 933 targets and 28 212 diseases to 7263 herbs and 49 258 ingredients and provided six pairwise relationships among them in HERB. In summary, HERB will intensively support the modernization of TCM and guide rational modern drug discovery efforts. And it is accessible through http://herb.ac.cn/.</p>]]></description>
            <pubDate><![CDATA[2020-12-02T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[MitImpact 3: modeling the residue interaction network of the Respiratory Chain subunits]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739348503-fca4d9d1-57d6-40c4-b68b-df5e6e2551d6/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1032</link>
            <description><![CDATA[<p class="para" id="N65541">Numerous lines of evidence have shown that the interaction between the nuclear and mitochondrial genomes ensures the efficient functioning of the OXPHOS complexes, with substantial implications in bioenergetics, adaptation, and disease. Their interaction is a fascinating and complex trait of the eukaryotic cell that MitImpact explores with its third major release. MitImpact expands its collection of genomic, clinical, and functional annotations of all non-synonymous substitutions of the human mitochondrial genome with new information on putative Compensated Pathogenic Deviations and co-varying amino acid sites of the Respiratory Chain subunits. It further provides evidence of energetic and structural residue compensation by techniques of molecular dynamics simulation. MitImpact is freely accessible at http://mitimpact.css-mendel.it.</p>]]></description>
            <pubDate><![CDATA[2020-12-09T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[MNDR v3.0: mammal ncRNA–disease repository with increased coverage and annotation]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739308713-24930ae4-af46-423e-8598-18452a41c575/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa707</link>
            <description><![CDATA[<p class="para" id="N65541">Many studies have indicated that non-coding RNA (ncRNA) dysfunction is closely related to numerous diseases. Recently, accumulated ncRNA–disease associations have made related databases insufficient to meet the demands of biomedical research. The constant updating of ncRNA–disease resources has become essential. Here, we have updated the mammal ncRNA–disease repository (MNDR, http://www.rna-society.org/mndr/) to version 3.0, containing more than one million entries, four-fold increment in data compared to the previous version. Experimental and predicted circRNA–disease associations have been integrated, increasing the number of categories of ncRNAs to five, and the number of mammalian species to 11. Moreover, ncRNA–disease related drug annotations and associations, as well as ncRNA subcellular localizations and interactions, were added. In addition, three ncRNA–disease (miRNA/lncRNA/circRNA) prediction tools were provided, and the website was also optimized, making it more practical and user-friendly. In summary, MNDR v3.0 will be a valuable resource for the investigation of disease mechanisms and clinical treatment strategies.</p>]]></description>
            <pubDate><![CDATA[2020-08-24T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The antiSMASH database version 3: increased taxonomic coverage and new query features for modular enzymes]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739296880-a22ea3c0-bf4a-4397-9c7b-d363b8d0ca11/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa978</link>
            <description><![CDATA[<p class="para" id="N65541">Microorganisms produce natural products that are frequently used in the development of antibacterial, antiviral, and anticancer drugs, pesticides, herbicides, or fungicides. In recent years, genome mining has evolved into a prominent method to access this potential. antiSMASH is one of the most popular tools for this task. Here, we present version 3 of the antiSMASH database, providing a means to access and query precomputed antiSMASH-5.2-detected biosynthetic gene clusters from representative, publicly available, high-quality microbial genomes via an interactive graphical user interface. In version 3, the database contains 147 517 high quality BGC regions from 388 archaeal, 25 236 bacterial and 177 fungal genomes and is available at https://antismash-db.secondarymetabolites.org/.</p><p class="para" id="N65542">
<div class="section" id="ga1"><div class="img"><div class="imgeVideo"><div class="img-fullscreenIcon" onClick="javascript:showImageContent('ga1');"><img src="/public/images/journalImg/fullscreen.png"/></div><div class="imageVideo"><img src="/dataresources/secured/content-1765739296880-a22ea3c0-bf4a-4397-9c7b-d363b8d0ca11/assets/gkaa978gra1.jpg" alt="The antiSMASH database was updated to version 3, now containing 147 517 secondary/specialized metabolite biosynthetic gene clusters (BGCs) originating from 25 236 bacterial, 388 archaeal and 177 fungal genomes. The database provides interactive access to the detailed annotation of each BGC that was generated with the most recent version of antiSMASH (v5.2)."/></div></div><div class="imgeVideoCaption" id="N65544"><div class="captionTitle">Graphical Abstract</div><div class="captionText">                                      The antiSMASH database was updated to version 3, now containing 147 517 secondary/specialized metabolite biosynthetic gene clusters (BGCs) originating from 25 236 bacterial, 388 archaeal and 177 fungal genomes. The database provides interactive access to the detailed annotation of each BGC that was generated with the most recent version of antiSMASH (v5.2).</div></div></div></div>
</p>]]></description>
            <pubDate><![CDATA[2020-11-05T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[jMorp updates in 2020: large enhancement of multi-omics data resources on the general Japanese population]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739258554-bd23256e-03cc-4c79-9e3f-6508e2a256a8/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1034</link>
            <description><![CDATA[<p class="para" id="N65541">In the Tohoku Medical Megabank project, genome and omics analyses of participants in two cohort studies were performed. A part of the data is available at the Japanese Multi Omics Reference Panel (jMorp; https://jmorp.megabank.tohoku.ac.jp) as a web-based database, as reported in our previous manuscript published in <i>Nucleic Acid Research</i> in 2018. At that time, jMorp mainly consisted of metabolome data; however, now genome, methylome, and transcriptome data have been integrated in addition to the enhancement of the number of samples for the metabolome data. For genomic data, jMorp provides a Japanese reference sequence obtained using <i>de novo</i> assembly of sequences from three Japanese individuals and allele frequencies obtained using whole-genome sequencing of 8,380 Japanese individuals. In addition, the omics data include methylome and transcriptome data from ∼300 samples and distribution of concentrations of more than 755 metabolites obtained using high-throughput nuclear magnetic resonance and high-sensitivity mass spectrometry. In summary, jMorp now provides four different kinds of omics data (genome, methylome, transcriptome, and metabolome), with a user-friendly web interface. This will be a useful scientific data resource on the general population for the discovery of disease biomarkers and personalized disease prevention and early diagnosis.</p>]]></description>
            <pubDate><![CDATA[2020-11-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Cyanorak v2.1: a scalable information system dedicated to the visualization and expert curation of marine and brackish picocyanobacteria genomes]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739253009-06830789-7270-45c1-9018-3024021c24ba/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa958</link>
            <description><![CDATA[<p class="para" id="N65541">Cyanorak v2.1 (http://www.sb-roscoff.fr/cyanorak) is an information system dedicated to visualizing, comparing and curating the genomes of <i>Prochlorococcus</i>, <i>Synechococcus</i> and <i>Cyanobium</i>, the most abundant photosynthetic microorganisms on Earth. The database encompasses sequences from 97 genomes, covering most of the wide genetic diversity known so far within these groups, and which were split into 25,834 clusters of likely orthologous groups (CLOGs). The user interface gives access to genomic characteristics, accession numbers as well as an interactive map showing strain isolation sites. The main entry to the database is through search for a term (gene name, product, etc.), resulting in a list of CLOGs and individual genes. Each CLOG benefits from a rich functional annotation including EggNOG, EC/K numbers, GO terms, TIGR Roles, custom-designed Cyanorak Roles as well as several protein motif predictions. Cyanorak also displays a phyletic profile, indicating the genotype and pigment type for each CLOG, and a genome viewer (Jbrowse) to visualize additional data on each genome such as predicted operons, genomic islands or transcriptomic data, when available. This information system also includes a BLAST search tool, comparative genomic context as well as various data export options. Altogether, Cyanorak v2.1 constitutes an invaluable, scalable tool for comparative genomics of ecologically relevant marine microorganisms.</p>]]></description>
            <pubDate><![CDATA[2020-10-30T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[ATACdb: a comprehensive human chromatin accessibility database]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739231057-bf00a877-a514-4019-8ff0-d7e681a3beab/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa943</link>
            <description><![CDATA[<p class="para" id="N65541">Accessible chromatin is a highly informative structural feature for identifying regulatory elements, which provides a large amount of information about transcriptional activity and gene regulatory mechanisms. Human ATAC-seq datasets are accumulating rapidly, prompting an urgent need to comprehensively collect and effectively process these data. We developed a comprehensive human chromatin accessibility database (ATACdb, http://www.licpathway.net/ATACdb), with the aim of providing a large amount of publicly available resources on human chromatin accessibility data, and to annotate and illustrate potential roles in a tissue/cell type-specific manner. The current version of ATACdb documented a total of 52 078 883 regions from over 1400 ATAC-seq samples. These samples have been manually curated from over 2200 chromatin accessibility samples from NCBI GEO/SRA. To make these datasets more accessible to the research community, ATACdb provides a quality assurance process including four quality control (QC) metrics. ATACdb provides detailed (epi)genetic annotations in chromatin accessibility regions, including super-enhancers, typical enhancers, transcription factors (TFs), common single-nucleotide polymorphisms (SNPs), risk SNPs, eQTLs, LD SNPs, methylations, chromatin interactions and TADs. Especially, ATACdb provides accurate inference of TF footprints within chromatin accessibility regions. ATACdb is a powerful platform that provides the most comprehensive accessible chromatin data, QC, TF footprint and various other annotations.</p>]]></description>
            <pubDate><![CDATA[2020-10-30T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[MASI: microbiota—active substance interactions database]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739193463-0f859d55-4c1c-4d8d-aafb-93384d3ce860/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa924</link>
            <description><![CDATA[<p class="para" id="N65541">Xenobiotic and host active substances interact with gut microbiota to influence human health and therapeutics. Dietary, pharmaceutical, herbal and environmental substances are modified by microbiota with altered bioavailabilities, bioactivities and toxic effects. Xenobiotics also affect microbiota with health implications. Knowledge of these microbiota and active substance interactions is important for understanding microbiota-regulated functions and therapeutics. Established microbiota databases provide useful information about the microbiota-disease associations, diet and drug interventions, and microbiota modulation of drugs. However, there is insufficient information on the active substances modified by microbiota and the abundance of gut bacteria in humans. Only ∼7% drugs are covered by the established databases. To complement these databases, we developed MASI, Microbiota—Active Substance Interactions database, for providing the information about the microbiota alteration of various substances, substance alteration of microbiota, and the abundance of gut bacteria in humans. These include 1,051 pharmaceutical, 103 dietary, 119 herbal, 46 probiotic, 142 environmental substances interacting with 806 microbiota species linked to 56 diseases and 784 microbiota–disease associations. MASI covers 11 215 bacteria-pharmaceutical, 914 bacteria-herbal, 309 bacteria-dietary, 753 bacteria-environmental substance interactions and the abundance profiles of 259 bacteria species in 3465 patients and 5334 healthy individuals. MASI is freely accessible at http://www.aiddlab.com/MASI.</p>]]></description>
            <pubDate><![CDATA[2020-10-30T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Peryton: a manual collection of experimentally supported microbe-disease associations]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739130693-ae6e05e3-0ed5-4d21-9424-d6e5a3fde073/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa902</link>
            <description><![CDATA[<p class="para" id="N65541">We present <i>Peryton</i> (https://dianalab.e-ce.uth.gr/peryton/), a database of experimentally supported microbe-disease associations. Its first version constitutes a novel resource hosting more than 7900 entries linking 43 diseases with 1396 microorganisms. <i>Peryton's</i> content is exclusively sustained by manual curation of biomedical articles. Diseases and microorganisms are provided in a systematic, standardized manner using reference resources to create database dictionaries. Information about the experimental design, study cohorts and the applied high- or low-throughput techniques is meticulously annotated and catered to users. Several functionalities are provided to enhance user experience and enable ingenious use of <i>Peryton</i>. One or more microorganisms and/or diseases can be queried at the same time. Advanced filtering options and direct text-based filtering of results enable refinement of returned information and the conducting of tailored queries suitable to different research questions. <i>Peryton</i> also provides interactive visualizations to effectively capture different aspects of its content and results can be directly downloaded for local storage and downstream analyses. <i>Peryton</i> will serve as a valuable source, enabling scientists of microbe-related disease fields to form novel hypotheses but, equally importantly, to assist in cross-validation of findings.</p>]]></description>
            <pubDate><![CDATA[2020-10-20T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[MethHC 2.0: information repository of DNA methylation and gene expression in human cancer]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765738992900-4666978d-c49b-4d55-83ea-69c13f4cc1bc/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1104</link>
            <description><![CDATA[<p class="para" id="N65541">DNA methylation is an important epigenetic regulator in gene expression and has several roles in cancer and disease progression. MethHC version 2.0 (MethHC 2.0) is an integrated and web-based resource focusing on the aberrant methylomes of human diseases, specifically cancer. This paper presents an updated implementation of MethHC 2.0 by incorporating additional DNA methylomes and transcriptomes from several public repositories, including 33 human cancers, over 50 118 microarray and RNA sequencing data from TCGA and GEO, and accumulating up to 3586 manually curated data from &gt;7000 collected published literature with experimental evidence. MethHC 2.0 has also been equipped with enhanced data annotation functionality and a user-friendly web interface for data presentation, search, and visualization. Provided features include clinical-pathological data, mutation and copy number variation, multiplicity of information (gene regions, enhancer regions, and CGI regions), and circulating tumor DNA methylation profiles, available for research such as biomarker panel design, cancer comparison, diagnosis, prognosis, therapy study and identifying potential epigenetic biomarkers. MethHC 2.0 is now available at http://awi.cuhk.edu.cn/∼MethHC.</p>]]></description>
            <pubDate><![CDATA[2020-12-03T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[PhycoCosm, a comparative algal genomics resource]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765738987679-d9903658-1d89-491b-87a8-faef9803f02e/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa898</link>
            <description><![CDATA[<p class="para" id="N65541">Algae are a diverse, polyphyletic group of photosynthetic eukaryotes spanning nearly all eukaryotic lineages of life and collectively responsible for ∼50% of photosynthesis on Earth. Sequenced algal genomes, critical to understanding their complex biology, are growing in number and require efficient tools for analysis. PhycoCosm (https://phycocosm.jgi.doe.gov) is an algal multi-omics portal, developed by the US Department of Energy Joint Genome Institute to support analysis and distribution of algal genome sequences and other ‘omics’ data. PhycoCosm provides integration of genome sequence and annotation for &gt;100 algal genomes with available multi-omics data and interactive web-based tools to enable algal research in bioenergy and the environment, encouraging community engagement and data exchange, and fostering new sequencing projects that will further these research goals.</p>]]></description>
            <pubDate><![CDATA[2020-10-26T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[OpenProt 2021: deeper functional annotation of the coding potential of eukaryotic genomes]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765738982746-9c6ca6da-c4aa-4516-b68f-f6665332aeb1/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1036</link>
            <description><![CDATA[<p class="para" id="N65541">OpenProt (www.openprot.org) is the first proteogenomic resource supporting a polycistronic annotation model for eukaryotic genomes. It provides a deeper annotation of open reading frames (ORFs) while mining experimental data for supporting evidence using cutting-edge algorithms. This update presents the major improvements since the initial release of OpenProt. All species support recent NCBI RefSeq and Ensembl annotations, with changes in annotations being reported in OpenProt. Using the 131 ribosome profiling datasets re-analysed by OpenProt to date, non-AUG initiation starts are reported alongside a confidence score of the initiating codon. From the 177 mass spectrometry datasets re-analysed by OpenProt to date, the unicity of the detected peptides is controlled at each implementation. Furthermore, to guide the users, detectability statistics and protein relationships (isoforms) are now reported for each protein. Finally, to foster access to deeper ORF annotation independently of one’s bioinformatics skills or computational resources, OpenProt now offers a data analysis platform. Users can submit their dataset for analysis and receive the results from the analysis by OpenProt. All data on OpenProt are freely available and downloadable for each species, the release-based format ensuring a continuous access to the data. Thus, OpenProt enables a more comprehensive annotation of eukaryotic genomes and fosters functional proteomic discoveries.</p>]]></description>
            <pubDate><![CDATA[2020-11-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[RNAcentral 2021: secondary structure integration, improved sequence search and new member databases]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765738976019-f26f43ef-8b7d-486f-bbd3-5d5983b86588/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa921</link>
            <description><![CDATA[<p class="para" id="N65541">RNAcentral is a comprehensive database of non-coding RNA (ncRNA) sequences that provides a single access point to 44 RNA resources and &gt;18 million ncRNA sequences from a wide range of organisms and RNA types. RNAcentral now also includes secondary (2D) structure information for &gt;13 million sequences, making RNAcentral the world’s largest RNA 2D structure database. The 2D diagrams are displayed using R2DT, a new 2D structure visualization method that uses consistent, reproducible and recognizable layouts for related RNAs. The sequence similarity search has been updated with a faster interface featuring facets for filtering search results by RNA type, organism, source database or any keyword. This sequence search tool is available as a reusable web component, and has been integrated into several RNAcentral member databases, including Rfam, miRBase and snoDB. To allow for a more fine-grained assignment of RNA types and subtypes, all RNAcentral sequences have been annotated with Sequence Ontology terms. The RNAcentral database continues to grow and provide a central data resource for the RNA community. RNAcentral is freely available at https://rnacentral.org.</p>]]></description>
            <pubDate><![CDATA[2020-10-27T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[crisprSQL: a novel database platform for CRISPR/Cas off-target cleavage assays]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610422100-4ed0ff1e-0ed8-41a4-add8-ab3bd08e54df/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa885</link>
            <description><![CDATA[<p class="para" id="N65541">With ongoing development of the CRISPR/Cas programmable nuclease system, applications in the area of <i>in vivo</i> therapeutic gene editing are increasingly within reach. However, non-negligible off-target effects remain a major concern for clinical applications. Even though a multitude of off-target cleavage datasets have been published, a comprehensive, transparent overview tool has not yet been established. Here, we present crisprSQL (http://www.crisprsql.com), an interactive and bioinformatically enhanced collection of CRISPR/Cas9 off-target cleavage studies aimed at enriching the fields of cleavage profiling, gene editing safety analysis and transcriptomics. The current version of crisprSQL contains cleavage data from 144 guide RNAs on 25,632 guide-target pairs from human and rodent cell lines, with interaction-specific references to epigenetic markers and gene names. The first curated database of this standard, it promises to enhance safety quantification research, inform experiment design and fuel development of computational off-target prediction algorithms.</p><p class="para" id="N65542">
<div class="section" id="ga1"><div class="img"><div class="imgeVideo"><div class="img-fullscreenIcon" onClick="javascript:showImageContent('ga1');"><img src="/public/images/journalImg/fullscreen.png"/></div><div class="imageVideo"><img src="/dataresources/secured/content-1765610422100-4ed0ff1e-0ed8-41a4-add8-ab3bd08e54df/assets/gkaa885gra1.jpg" alt="Overview of the crisprSQL database. Cleavage and epigenetics data (left) are joined together and supplemented by attributes calculated from their properties (top). This forms the crisprSQL database, which has an online user interface supporting data browsing, export and upload of new studies (right)."/></div></div><div class="imgeVideoCaption" id="N65544"><div class="captionTitle">Graphical Abstract</div><div class="captionText">                                      Overview of the crisprSQL database. Cleavage and epigenetics data (left) are joined together and supplemented by attributes calculated from their properties (top). This forms the crisprSQL database, which has an online user interface supporting data browsing, export and upload of new studies (right).</div></div></div></div>
</p>]]></description>
            <pubDate><![CDATA[2020-10-21T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[LncSEA: a platform for long non-coding RNA related sets and enrichment analysis]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610375866-8e1fa084-ee59-4b0f-8473-f4e510039c55/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa806</link>
            <description><![CDATA[<p class="para" id="N65541">Long non-coding RNAs (lncRNAs) have been proven to play important roles in transcriptional processes and various biological functions. Establishing a comprehensive collection of human lncRNA sets is urgent work at present. Using reference lncRNA sets, enrichment analyses will be useful for analyzing lncRNA lists of interest submitted by users. Therefore, we developed a human lncRNA sets database, called LncSEA, which aimed to document a large number of available resources for human lncRNA sets and provide annotation and enrichment analyses for lncRNAs. LncSEA supports &gt;40 000 lncRNA reference sets across 18 categories and 66 sub-categories, and covers over 50 000 lncRNAs. We not only collected lncRNA sets based on downstream regulatory data sources, but also identified a large number of lncRNA sets regulated by upstream transcription factors (TFs) and DNA regulatory elements by integrating TF ChIP-seq, DNase-seq, ATAC-seq and H3K27ac ChIP-seq data. Importantly, LncSEA provides annotation and enrichment analyses of lncRNA sets associated with upstream regulators and downstream targets. In summary, LncSEA is a powerful platform that provides a variety of types of lncRNA sets for users, and supports lncRNA annotations and enrichment analyses. The LncSEA database is freely accessible at http://bio.liclab.net/LncSEA/index.php.</p>]]></description>
            <pubDate><![CDATA[2020-10-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[canSAR: update to the cancer translational research and drug discovery knowledgebase]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610371258-1bcf18c3-9792-4a6b-b919-cf76c1df1a99/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1059</link>
            <description><![CDATA[<p class="para" id="N65541">canSAR (http://cansar.icr.ac.uk) is the largest, public, freely available, integrative translational research and drug discovery knowledgebase for oncology. canSAR integrates vast multidisciplinary data from across genomic, protein, pharmacological, drug and chemical data with structural biology, protein networks and more. It also provides unique data, curation and annotation and crucially, AI-informed target assessment for drug discovery. canSAR is widely used internationally by academia and industry. Here we describe significant developments and enhancements to the data, web interface and infrastructure of canSAR in the form of the new implementation of the system: canSAR<i>black</i>. We demonstrate new functionality in aiding translation hypothesis generation and experimental design, and show how canSAR can be adapted and utilised outside oncology.</p>]]></description>
            <pubDate><![CDATA[2020-11-21T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[REDIportal: millions of novel A-to-I RNA editing events from thousands of RNAseq experiments]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610307533-7ddd1c4a-0ae7-4e99-a8d0-6df6383c5049/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa916</link>
            <description><![CDATA[<p class="para" id="N65541">RNA editing is a relevant epitranscriptome phenomenon able to increase the transcriptome and proteome diversity of eukaryotic organisms. ADAR mediated RNA editing is widespread in humans in which millions of A-to-I changes modify thousands of primary transcripts. RNA editing has pivotal roles in the regulation of gene expression or modulation of the innate immune response or functioning of several neurotransmitter receptors. Massive transcriptome sequencing has fostered the research in this field. Nonetheless, different aspects of the RNA editing biology are still unknown and need to be elucidated. To support the study of A-to-I RNA editing we have updated our REDIportal catalogue raising its content to about 16 millions of events detected in 9642 human RNAseq samples from the GTEx project by using a dedicated pipeline based on the HPC version of the REDItools software. REDIportal now allows searches at sample level, provides overviews of RNA editing profiles per each RNAseq experiment, implements a Gene View module to look at individual events in their genic context and hosts the CLAIRE database. Starting from this novel version, REDIportal will start collecting non-human RNA editing changes for comparative genomics investigations. The database is freely available at http://srv00.recas.ba.infn.it/atlas/index.html.</p>]]></description>
            <pubDate><![CDATA[2020-10-26T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[miRNASNP-v3: a comprehensive database for SNPs and disease-related variations in miRNAs and miRNA targets]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610303771-24572ff6-7072-4468-a72a-6f82b832b33e/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa783</link>
            <description><![CDATA[<p class="para" id="N65541">MicroRNAs (miRNAs) related single-nucleotide variations (SNVs), including single-nucleotide polymorphisms (SNPs) and disease-related variations (DRVs) in miRNAs and miRNA-target binding sites, can affect miRNA functions and/or biogenesis, thus to impact on phenotypes. miRNASNP is a widely used database for miRNA-related SNPs and their effects. Here, we updated it to miRNASNP-v3 (http://bioinfo.life.hust.edu.cn/miRNASNP/) with tremendous number of SNVs and new features, especially the DRVs data. We analyzed the effects of 7 161 741 SNPs and 505 417 DRVs on 1897 pre-miRNAs (2630 mature miRNAs) and 3′UTRs of 18 152 genes. miRNASNP-v3 provides a one-stop resource for miRNA-related SNVs research with the following functions: (i) explore associations between miRNA-related SNPs/DRVs and diseases; (ii) browse the effects of SNPs/DRVs on miRNA-target binding; (iii) functional enrichment analysis of miRNA target gain/loss caused by SNPs/DRVs; (iv) investigate correlations between drug sensitivity and miRNA expression; (v) inquire expression profiles of miRNAs and their targets in cancers; (vi) browse the effects of SNPs/DRVs on pre-miRNA secondary structure changes; and (vii) predict the effects of user-defined variations on miRNA-target binding or pre-miRNA secondary structure. miRNASNP-v3 is a valuable and long-term supported resource in functional variation screening and miRNA function studies.</p>]]></description>
            <pubDate><![CDATA[2020-09-29T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[HRT Atlas v1.0 database: redefining human and mouse housekeeping genes and candidate reference transcripts by mining massive RNA-seq datasets]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610290670-0da88ee0-0fcf-48a9-9fc5-3f946cad4b68/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa609</link>
            <description><![CDATA[<p class="para" id="N65541">Housekeeping (HK) genes are constitutively expressed genes that are required for the maintenance of basic cellular functions. Despite their importance in the calibration of gene expression, as well as the understanding of many genomic and evolutionary features, important discrepancies have been observed in studies that previously identified these genes. Here, we present Housekeeping and Reference Transcript Atlas (HRT Atlas v1.0, www.housekeeping.unicamp.br) a web-based database which addresses some of the previously observed limitations in the identification of these genes, and offers a more accurate database of human and mouse HK genes and transcripts. The database was generated by mining massive human and mouse RNA-seq data sets, including 11 281 and 507 high-quality RNA-seq samples from 52 human non-disease tissues/cells and 14 healthy tissues/cells of C57BL/6 wild type mouse, respectively. User can visualize the expression and download lists of 2158 human HK transcripts from 2176 HK genes and 3024 mouse HK transcripts from 3277 mouse HK genes. HRT Atlas also offers the most stable and suitable tissue selective candidate reference transcripts for normalization of qPCR experiments. Specific primers and predicted modifiers of gene expression for some of these HK transcripts are also proposed. HRT Atlas has also been integrated with a regulatory elements resource from Epiregio server.</p><p class="para" id="N65542">
<div class="section" id="ga1"><div class="img"><div class="imgeVideo"><div class="img-fullscreenIcon" onClick="javascript:showImageContent('ga1');"><img src="/public/images/journalImg/fullscreen.png"/></div><div class="imageVideo"><img src="/dataresources/secured/content-1765610290670-0da88ee0-0fcf-48a9-9fc5-3f946cad4b68/assets/gkaa609gra1.jpg" alt="The workflow of database generation and web-tool implementation are shown. After downloading and identification of HK genes and transcripts using a specific algorithm, transcripts of genes with unknown pseudogenes were used to select suitable candidate reference transcripts. Some candidate reference transcript-specific primers were designed. The web-based tool was developed using Shiny package and encapsulated into docker image for deployment."/></div></div><div class="imgeVideoCaption" id="N65544"><div class="captionTitle">Graphical Abstract</div><div class="captionText">                                      The workflow of database generation and web-tool implementation are shown. After downloading and identification of HK genes and transcripts using a specific algorithm, transcripts of genes with unknown pseudogenes were used to select suitable candidate reference transcripts. Some candidate reference transcript-specific primers were designed. The web-based tool was developed using Shiny package and encapsulated into docker image for deployment.</div></div></div></div>
</p>]]></description>
            <pubDate><![CDATA[2020-07-14T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The Bgee suite: integrated curated expression atlas and comparative transcriptomics in animals]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610255022-32a1a5a5-2962-4fc6-99cf-82d9d7475b4b/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa793</link>
            <description><![CDATA[<p class="para" id="N65541">Bgee is a database to retrieve and compare gene expression patterns in multiple animal species, produced by integrating multiple data types (RNA-Seq, Affymetrix, in situ hybridization, and EST data). It is based exclusively on curated healthy wild-type expression data (e.g., no gene knock-out, no treatment, no disease), to provide a comparable reference of normal gene expression. Curation includes very large datasets such as GTEx (re-annotation of samples as ‘healthy’ or not) as well as many small ones. Data are integrated and made comparable between species thanks to consistent data annotation and processing, and to calls of presence/absence of expression, along with expression scores. As a result, Bgee is capable of detecting the conditions of expression of any single gene, accommodating any data type and species. Bgee provides several tools for analyses, allowing, e.g., automated comparisons of gene expression patterns within and between species, retrieval of the prefered conditions of expression of any gene, or enrichment analyses of conditions with expression of sets of genes. Bgee release 14.1 includes 29 animal species, and is available at https://bgee.org/ and through its Bioconductor R package BgeeDB.</p>]]></description>
            <pubDate><![CDATA[2020-10-10T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[HumanMetagenomeDB: a public repository of curated and standardized metadata for human metagenomes]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610227224-9b78032e-c77c-4bf3-bfd8-6be8585e6684/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1031</link>
            <description><![CDATA[<p class="para" id="N65541">Metagenomics became a standard strategy to comprehend the functional potential of microbial communities, including the human microbiome. Currently, the number of metagenomes in public repositories is increasing exponentially. The Sequence Read Archive (SRA) and the MG-RAST are the two main repositories for metagenomic data. These databases allow scientists to reanalyze samples and explore new hypotheses. However, mining samples from them can be a limiting factor, since the metadata available in these repositories is often misannotated, misleading, and decentralized, creating an overly complex environment for sample reanalysis. The main goal of the HumanMetagenomeDB is to simplify the identification and use of public human metagenomes of interest. HumanMetagenomeDB version 1.0 contains metadata of 69 822 metagenomes. We standardized 203 attributes, based on standardized ontologies, describing host characteristics (e.g. sex, age and body mass index), diagnosis information (e.g. cancer, Crohn's disease and Parkinson), location (e.g. country, longitude and latitude), sampling site (e.g. gut, lung and skin) and sequencing attributes (e.g. sequencing platform, average length and sequence quality). Further, HumanMetagenomeDB version 1.0 metagenomes encompass 58 countries, 9 main sample sites (i.e. body parts), 58 diagnoses and multiple ages, ranging from just born to 91 years old. The HumanMetagenomeDB is publicly available at https://webapp.ufz.de/hmgdb/.</p>]]></description>
            <pubDate><![CDATA[2020-11-22T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[MitoCarta3.0: an updated mitochondrial proteome now with sub-organelle localization and pathway annotations]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610211924-cf4ae0f7-7c5f-46f3-83cc-d302cf814523/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1011</link>
            <description><![CDATA[<p class="para" id="N65541">The mammalian mitochondrial proteome is under dual genomic control, with 99% of proteins encoded by the nuclear genome and 13 originating from the mitochondrial DNA (mtDNA). We previously developed MitoCarta, a catalogue of over 1000 genes encoding the mammalian mitochondrial proteome. This catalogue was compiled using a Bayesian integration of multiple sequence features and experimental datasets, notably protein mass spectrometry of mitochondria isolated from fourteen murine tissues. Here, we introduce MitoCarta3.0. Beginning with the MitoCarta2.0 inventory, we performed manual review to remove 100 genes and introduce 78 additional genes, arriving at an updated inventory of 1136 human genes. We now include manually curated annotations of sub-mitochondrial localization (matrix, inner membrane, intermembrane space, outer membrane) as well as assignment to 149 hierarchical ‘MitoPathways’ spanning seven broad functional categories relevant to mitochondria. MitoCarta3.0, including sub-mitochondrial localization and MitoPathway annotations, is freely available at http://www.broadinstitute.org/mitocarta and should serve as a continued community resource for mitochondrial biology and medicine.</p>]]></description>
            <pubDate><![CDATA[2020-11-11T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[MoonProt 3.0: an update of the moonlighting proteins database]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610208202-aba37c73-370d-45a3-9f42-d09984ac0e39/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1101</link>
            <description><![CDATA[<p class="para" id="N65541">MoonProt 3.0 (http://moonlightingproteins.org) is an updated open-access database storing expert-curated annotations for moonlighting proteins. Moonlighting proteins have two or more physiologically relevant distinct biochemical or biophysical functions performed by a single polypeptide chain. Here, we describe an expansion in the database since our previous report in the Database Issue of Nucleic Acids Research in 2018. For this release, the number of proteins annotated has been expanded to over 500 proteins and dozens of protein annotations have been updated with additional information, including more structures in the Protein Data Bank, compared with version 2.0. The new entries include more examples from humans, plants and archaea, more proteins involved in disease and proteins with different combinations of functions. More kinds of information about the proteins and the species in which they have multiple functions has been added, including CATH and SCOP classification of structure, known and predicted disorder, predicted transmembrane helices, type of organism, relationship of the protein to disease, and relationship of organism to cause of disease.</p>]]></description>
            <pubDate><![CDATA[2020-11-27T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[GENCODE 2021]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610195762-a35e1146-f31a-4d3d-a00d-d7c0eb834cb8/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1087</link>
            <description><![CDATA[<p class="para" id="N65541">The GENCODE project annotates human and mouse genes and transcripts supported by experimental data with high accuracy, providing a foundational resource that supports genome biology and clinical genomics. GENCODE annotation processes make use of primary data and bioinformatic tools and analysis generated both within the consortium and externally to support the creation of transcript structures and the determination of their function. Here, we present improvements to our annotation infrastructure, bioinformatics tools, and analysis, and the advances they support in the annotation of the human and mouse genomes including: the completion of first pass manual annotation for the mouse reference genome; targeted improvements to the annotation of genes associated with SARS-CoV-2 infection; collaborative projects to achieve convergence across reference annotation databases for the annotation of human and mouse protein-coding genes; and the first GENCODE manually supervised automated annotation of lncRNAs. Our annotation is accessible via Ensembl, the UCSC Genome Browser and https://www.gencodegenes.org.</p>]]></description>
            <pubDate><![CDATA[2020-12-03T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[RCSB Protein Data Bank: powerful new tools for exploring 3D structures of biological macromolecules for basic and applied research and education in fundamental biology, biomedicine, biotechnology, bioengineering and energy sciences]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610188814-cf50180a-949c-4902-ad34-b98778ad48c3/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1038</link>
            <description><![CDATA[<p class="para" id="N65541">The Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB), the US data center for the global PDB archive and a founding member of the Worldwide Protein Data Bank partnership, serves tens of thousands of data depositors in the Americas and Oceania and makes 3D macromolecular structure data available at no charge and without restrictions to millions of RCSB.org users around the world, including &gt;660 000 educators, students and members of the curious public using PDB101.RCSB.org. PDB data depositors include structural biologists using macromolecular crystallography, nuclear magnetic resonance spectroscopy, 3D electron microscopy and micro-electron diffraction. PDB data consumers accessing our web portals include researchers, educators and students studying fundamental biology, biomedicine, biotechnology, bioengineering and energy sciences. During the past 2 years, the research-focused RCSB PDB web portal (RCSB.org) has undergone a complete redesign, enabling improved searching with full Boolean operator logic and more facile access to PDB data integrated with &gt;40 external biodata resources. New features and resources are described in detail using examples that showcase recently released structures of SARS-CoV-2 proteins and host cell proteins relevant to understanding and addressing the COVID-19 global pandemic.</p>]]></description>
            <pubDate><![CDATA[2020-11-19T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[SilencerDB: a comprehensive database of silencers]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610142577-051bba41-4c3e-4f0f-81dc-7ca343842485/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa839</link>
            <description><![CDATA[<p class="para" id="N65541">Gene regulatory elements, including promoters, enhancers, silencers, etc., control transcriptional programs in a spatiotemporal manner. Though these elements are known to be able to induce either positive or negative transcriptional control, the community has been mostly studying enhancers which amplify transcription initiation, with less emphasis given to silencers which repress gene expression. To facilitate the study of silencers and the investigation of their potential roles in transcriptional control, we developed SilencerDB (http://health.tsinghua.edu.cn/silencerdb/), a comprehensive database of silencers by manually curating silencers from 2300 published articles. The current version, SilencerDB 1.0, contains (1) 33 060 validated silencers from experimental methods, and (ii) 5 045 547 predicted silencers from state-of-the-art machine learning methods. The functionality of SilencerDB includes (a) standardized categorization of silencers in a tree-structured class hierarchy based on species, organ, tissue and cell line and (b) comprehensive annotations of silencers with the nearest gene and potential regulatory genes. SilencerDB, to the best of our knowledge, is the first comprehensive database at this scale dedicated to silencers, with reliable annotations and user-friendly interactive database features. We believe this database has the potential to enable advanced understanding of silencers in regulatory mechanisms and to empower researchers to devise diverse applications of silencers in disease development.</p>]]></description>
            <pubDate><![CDATA[2020-10-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[VARAdb: a comprehensive variation annotation database for human]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610111778-dcc838fd-16e7-492a-ad84-79d93f94cd18/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa922</link>
            <description><![CDATA[<p class="para" id="N65541">With the study of human diseases and biological processes increasing, a large number of non-coding variants have been identified and facilitated. The rapid accumulation of genetic and epigenomic information has resulted in an urgent need to collect and process data to explore the regulation of non-coding variants. Here, we developed a comprehensive variation annotation database for human (VARAdb, http://www.licpathway.net/VARAdb/), which specifically considers non-coding variants. VARAdb provides annotation information for 577,283,813 variations and novel variants, prioritizes variations based on scores using nine annotation categories, and supports pathway downstream analysis. Importantly, VARAdb integrates a large amount of genetic and epigenomic data into five annotation sections, which include ‘Variation information’, ‘Regulatory information’, ‘Related genes’, ‘Chromatin accessibility’ and ‘Chromatin interaction’. The detailed annotation information consists of motif changes, risk SNPs, LD SNPs, eQTLs, clinical variant-drug-gene pairs, sequence conservation, somatic mutations, enhancers, super enhancers, promoters, transcription factors, chromatin states, histone modifications, chromatin accessibility regions and chromatin interactions. This database is a user-friendly interface to query, browse and visualize variations and related annotation information. VARAdb is a useful resource for selecting potential functional variations and interpreting their effects on human diseases and biological processes.</p>]]></description>
            <pubDate><![CDATA[2020-10-23T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Open Targets Genetics: systematic identification of trait-associated genes using large-scale genetics and functional genomics]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610097728-15731f9b-869e-401d-83cd-2997c6612370/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa840</link>
            <description><![CDATA[<p class="para" id="N65541">Open Targets Genetics (https://genetics.opentargets.org) is an open-access integrative resource that aggregates human GWAS and functional genomics data including gene expression, protein abundance, chromatin interaction and conformation data from a wide range of cell types and tissues to make robust connections between GWAS-associated loci, variants and likely causal genes. This enables systematic identification and prioritisation of likely causal variants and genes across all published trait-associated loci. In this paper, we describe the public resources we aggregate, the technology and analyses we use, and the functionality that the portal offers. Open Targets Genetics can be searched by variant, gene or study/phenotype. It offers tools that enable users to prioritise causal variants and genes at disease-associated loci and access systematic cross-disease and disease-molecular trait colocalization analysis across 92 cell types and tissues including the eQTL Catalogue. Data visualizations such as Manhattan-like plots, regional plots, credible sets overlap between studies and PheWAS plots enable users to explore GWAS signals in depth. The integrated data is made available through the web portal, for bulk download and via a GraphQL API, and the software is open source. Applications of this integrated data include identification of novel targets for drug discovery and drug repurposing.</p>]]></description>
            <pubDate><![CDATA[2020-10-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[KLIFS: an overhaul after the first 5 years of supporting kinase research]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610088118-f365c2ce-423d-456d-8d4b-fac8f5a8949d/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa895</link>
            <description><![CDATA[<p class="para" id="N65541">Kinases are a prime target of drug development efforts with &gt;60 drug approvals in the past two decades. Due to the research into this protein family, a wealth of data has been accumulated that keeps on growing. KLIFS—Kinase–Ligand Interaction Fingerprints and Structures—is a structural database focusing on how kinase inhibitors interact with their targets. The aim of KLIFS is to support (structure-based) kinase research through the systematic collection, annotation, and processing of kinase structures. Now, 5 years after releasing the initial KLIFS website, the database has undergone a complete overhaul with a new website, new logo, and new functionalities. In this article, we start by looking back at how KLIFS has been used by the research community, followed by a description of the renewed KLIFS, and conclude with showcasing the functionalities of KLIFS. Major changes include the integration of approved drugs and inhibitors in clinical trials, extension of the coverage to atypical kinases, and a RESTful API for programmatic access. KLIFS is available at the new domain https://klifs.net.</p>]]></description>
            <pubDate><![CDATA[2020-10-21T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[GPCRdb in 2021: integrating GPCR sequence, structure and function]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610051912-b6f4497e-643f-423f-aefc-80698068acbe/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1080</link>
            <description><![CDATA[<p class="para" id="N65541">G protein-coupled receptors (GPCRs) form both the largest family of membrane proteins and drug targets, mediating the action of one-third of medicines. The GPCR database, GPCRdb serves &gt;4 000 researchers every month and offers reference data, analysis of own or literature data, experiment design and dissemination of published datasets. Here, we describe new and updated GPCRdb resources with a particular focus on integration of sequence, structure and function. GPCRdb contains all human non-olfactory GPCRs (and &gt;27 000 orthologs), G-proteins and arrestins. It includes over 2 000 drug and in-trial agents and nearly 200 000 ligands with activity and availability data. GPCRdb annotates all published GPCR structures (updated monthly), which are also offered in a refined version (with re-modeled missing/distorted regions and reverted mutations) and provides structure models of all human non-olfactory receptors in inactive, intermediate and active states. Mutagenesis data in the GPCRdb spans natural genetic variants, GPCR-G protein interfaces, ligand sites and thermostabilising mutations. A new sequence signature tool for identification of functional residue determinants has been added and two data driven tools to design ligand site mutations and constructs for structure determination have been updated extending their coverage of receptors and modifications. The GPCRdb is available at https://gpcrdb.org.</p>]]></description>
            <pubDate><![CDATA[2020-12-03T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[PubChem in 2021: new data content and improved web interfaces]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610047183-cca649f1-14bb-4994-8716-975daa829e27/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa971</link>
            <description><![CDATA[<p class="para" id="N65541">PubChem (https://pubchem.ncbi.nlm.nih.gov) is a popular chemical information resource that serves the scientific community as well as the general public, with millions of unique users per month. In the past two years, PubChem made substantial improvements. Data from more than 100 new data sources were added to PubChem, including chemical-literature links from Thieme Chemistry, chemical and physical property links from SpringerMaterials, and patent links from the World Intellectual Properties Organization (WIPO). PubChem's homepage and individual record pages were updated to help users find desired information faster. This update involved a data model change for the data objects used by these pages as well as by programmatic users. Several new services were introduced, including the PubChem Periodic Table and Element pages, Pathway pages, and Knowledge panels. Additionally, in response to the coronavirus disease 2019 (COVID-19) outbreak, PubChem created a special data collection that contains PubChem data related to COVID-19 and the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2).</p>]]></description>
            <pubDate><![CDATA[2020-11-05T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[gcType: a high-quality type strain genome database for microbial phylogenetic and functional research]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610040372-d5e5abf9-e647-4ddf-b0b6-f0a58261b8b3/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa957</link>
            <description><![CDATA[<p class="para" id="N65541">Taxonomic and functional research of microorganisms has increasingly relied upon genome-based data and methods. As the depository of the Global Catalogue of Microorganisms (GCM) 10K prokaryotic type strain sequencing project, Global Catalogue of Type Strain (gcType) has published 1049 type strain genomes sequenced by the GCM 10K project which are preserved in global culture collections with a valid published status. Additionally, the information provided through gcType includes &gt;12 000 publicly available type strain genome sequences from GenBank incorporated using quality control criteria and standard data annotation pipelines to form a high-quality reference database. This database integrates type strain sequences with their phenotypic information to facilitate phenotypic and genotypic analyses. Multiple formats of cross-genome searches and interactive interfaces have allowed extensive exploration of the database's resources. In this study, we describe web-based data analysis pipelines for genomic analyses and genome-based taxonomy, which could serve as a one-stop platform for the identification of prokaryotic species. The number of type strain genomes that are published will continue to increase as the GCM 10K project increases its collaboration with culture collections worldwide. Data of this project is shared with the International Nucleotide Sequence Database Collaboration. Access to gcType is free at http://gctype.wdcm.org/.</p>]]></description>
            <pubDate><![CDATA[2020-10-29T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[ThermoMutDB: a thermodynamic database for missense mutations]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610035939-38ef19f9-212d-4432-acf8-7b82a4e2bf0e/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa925</link>
            <description><![CDATA[<p class="para" id="N65541">Proteins are intricate, dynamic structures, and small changes in their amino acid sequences can lead to large effects on their folding, stability and dynamics. To facilitate the further development and evaluation of methods to predict these changes, we have developed ThermoMutDB, a manually curated database containing &gt;14,669 experimental data of thermodynamic parameters for wild type and mutant proteins. This represents an increase of 83% in unique mutations over previous databases and includes thermodynamic information on 204 new proteins. During manual curation we have also corrected annotation errors in previously curated entries. Associated with each entry, we have included information on the unfolding Gibbs free energy and melting temperature change, and have associated entries with available experimental structural information. ThermoMutDB supports users to contribute to new data points and programmatic access to the database via a RESTful API. ThermoMutDB is freely available at: http://biosig.unimelb.edu.au/thermomutdb.</p><p class="para" id="N65542">
<div class="section" id="ga1"><div class="img"><div class="imgeVideo"><div class="img-fullscreenIcon" onClick="javascript:showImageContent('ga1');"><img src="/public/images/journalImg/fullscreen.png"/></div><div class="imageVideo"><img src="/dataresources/secured/content-1765610035939-38ef19f9-212d-4432-acf8-7b82a4e2bf0e/assets/gkaa925gra1.jpg" alt="ThermoMutDB is an online resource associating effects of missense mutations on protein thermodynamics."/></div></div><div class="imgeVideoCaption" id="N65544"><div class="captionTitle">Graphical Abstract</div><div class="captionText">                                      ThermoMutDB is an online resource associating effects of missense mutations on protein thermodynamics.</div></div></div></div>
</p>]]></description>
            <pubDate><![CDATA[2020-10-23T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The Human Phenotype Ontology in 2021]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610019914-76affff9-255a-411c-baf2-3a58299300c8/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1043</link>
            <description><![CDATA[<p class="para" id="N65541">The Human Phenotype Ontology (HPO, https://hpo.jax.org) was launched in 2008 to provide a comprehensive logical standard to describe and computationally analyze phenotypic abnormalities found in human disease. The HPO is now a worldwide standard for phenotype exchange. The HPO has grown steadily since its inception due to considerable contributions from clinical experts and researchers from a diverse range of disciplines. Here, we present recent major extensions of the HPO for neurology, nephrology, immunology, pulmonology, newborn screening, and other areas. For example, the seizure subontology now reflects the International League Against Epilepsy (ILAE) guidelines and these enhancements have already shown clinical validity. We present new efforts to harmonize computational definitions of phenotypic abnormalities across the HPO and multiple phenotype ontologies used for animal models of disease. These efforts will benefit software such as Exomiser by improving the accuracy and scope of cross-species phenotype matching. The computational modeling strategy used by the HPO to define disease entities and phenotypic features and distinguish between them is explained in detail.We also report on recent efforts to translate the HPO into indigenous languages. Finally, we summarize recent advances in the use of HPO in electronic health record systems.</p>]]></description>
            <pubDate><![CDATA[2020-12-02T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[AtMAD: <i>Arabidopsis thaliana</i> multi-omics association database]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610010109-7150d2c2-3218-4868-9ea5-1887c98b8c40/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1042</link>
            <description><![CDATA[<p class="para" id="N65541">Integration analysis of multi-omics data provides a comprehensive landscape for understanding biological systems and mechanisms. The abundance of high-quality multi-omics data (genomics, transcriptomics, methylomics and phenomics) for the model organism <i>Arabidopsis thaliana</i> enables scientists to study the genetic mechanism of many biological processes. However, no resource is available to provide comprehensive and systematic multi-omics associations for Arabidopsis. Here, we developed an <i>Arabidopsis thaliana</i> Multi-omics Association Database (AtMAD, http://www.megabionet.org/atmad), a public repository for large-scale measurements of associations between genome, transcriptome, methylome, pathway and phenotype in Arabidopsis, designed for facilitating identification of eQTL, emQTL, Pathway-mQTL, Phenotype-pathway, GWAS, TWAS and EWAS. Candidate variants/methylations/genes were identified in AtMAD for specific phenotypes or biological processes, many of them are supported by experimental evidence. Based on the multi-omics association strategy, we have identified 11 796 <i>cis</i>-eQTLs and 10 119 <i>trans</i>-eQTLs. Among them, 68 837 environment-eQTL associations and 149 622 GWAS-eQTL associations were identified and stored in AtMAD. For expression–methylation quantitative trait loci (emQTL), we identified 265 776 emQTLs and 122 344 pathway-mQTLs. For TWAS and EWAS, we obtained 62 754 significant phenotype-gene associations and 3 993 379 significant phenotype-methylation associations, respectively. Overall, the multi-omics associated network in AtMAD will provide new insights into exploring biological mechanisms of plants at multi-omics levels.</p>]]></description>
            <pubDate><![CDATA[2020-11-21T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[ViruSurf: an integrated database to investigate viral sequences]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609999557-358d7fb6-ad6b-4dc9-965a-b7db8a64f0d3/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa846</link>
            <description><![CDATA[<p class="para" id="N65541">ViruSurf, available at http://gmql.eu/virusurf/, is a large public database of viral sequences and integrated and curated metadata from heterogeneous sources (RefSeq, GenBank, COG-UK and NMDC); it also exposes computed nucleotide and amino acid variants, called from original sequences. A GISAID-specific ViruSurf database, available at http://gmql.eu/virusurf_gisaid/, offers a subset of these functionalities. Given the current pandemic outbreak, SARS-CoV-2 data are collected from the four sources; but ViruSurf contains other virus species harmful to humans, including SARS-CoV, MERS-CoV, Ebola and Dengue. The database is centered on sequences, described from their biological, technological and organizational dimensions. In addition, the analytical dimension characterizes the sequence in terms of its annotations and variants. The web interface enables expressing complex search queries in a simple way; arbitrary search queries can freely combine conditions on attributes from the four dimensions, extracting the resulting sequences. Several example queries on the database confirm and possibly improve results from recent research papers; results can be recomputed over time and upon selected populations. Effective search over large and curated sequence data may enable faster responses to future threats that could arise from new viruses.</p>]]></description>
            <pubDate><![CDATA[2020-10-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The Nucleome Data Bank: web-based resources to simulate and analyze the three-dimensional genome]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609994748-d83bbac7-ca33-4353-9f28-f50c2452047d/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa818</link>
            <description><![CDATA[<p class="para" id="N65541">We introduce the Nucleome Data Bank (NDB), a web-based platform to simulate and analyze the three-dimensional (3D) organization of genomes. The NDB enables physics-based simulation of chromosomal structural dynamics through the MEGABASE + MiChroM computational pipeline. The input of the pipeline consists of epigenetic information sourced from the Encode database; the output consists of the trajectories of chromosomal motions that accurately predict Hi-C and fluorescence <i>in</i><i>situ</i> hybridization data, as well as multiple observations of chromosomal dynamics <i>in vivo</i>. As an intermediate step, users can also generate chromosomal sub-compartment annotations directly from the same epigenetic input, without the use of any DNA–DNA proximity ligation data. Additionally, the NDB freely hosts both experimental and computational structural genomics data. Besides being able to perform their own genome simulations and download the hosted data, users can also analyze and visualize the same data through custom-designed web-based tools. In particular, the one-dimensional genetic and epigenetic data can be overlaid onto accurate 3D structures of chromosomes, to study the spatial distribution of genetic and epigenetic features. The NDB aims to be a shared resource to biologists, biophysicists and all genome scientists. The NDB is available at https://ndb.rice.edu.</p>]]></description>
            <pubDate><![CDATA[2020-10-06T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Pfam: The protein families database in 2021]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609980507-8d13f41c-e93c-415f-817c-3b6bb2aa933e/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa913</link>
            <description><![CDATA[<p class="para" id="N65541">The Pfam database is a widely used resource for classifying protein sequences into families and domains. Since Pfam was last described in this journal, over 350 new families have been added in Pfam 33.1 and numerous improvements have been made to existing entries. To facilitate research on COVID-19, we have revised the Pfam entries that cover the SARS-CoV-2 proteome, and built new entries for regions that were not covered by Pfam. We have reintroduced Pfam-B which provides an automatically generated supplement to Pfam and contains 136 730 novel clusters of sequences that are not yet matched by a Pfam family. The new Pfam-B is based on a clustering by the MMseqs2 software. We have compared all of the regions in the RepeatsDB to those in Pfam and have started to use the results to build and refine Pfam repeat families. Pfam is freely available for browsing and download at http://pfam.xfam.org/.</p>]]></description>
            <pubDate><![CDATA[2020-10-30T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Project Score database: a resource for investigating cancer cell dependencies and prioritizing therapeutic targets]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609976563-df99495a-ce68-41e4-81dd-6ec749a272a9/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa882</link>
            <description><![CDATA[<p class="para" id="N65541">CRISPR genetic screens in cancer cell models are a powerful tool to elucidate oncogenic mechanisms and to identify promising therapeutic targets. The Project Score database (https://score.depmap.sanger.ac.uk/) uses genome-wide CRISPR–Cas9 dropout screening data in hundreds of highly annotated cancer cell models to identify genes required for cell fitness and prioritize novel oncology targets. The Project Score database currently allows users to investigate the fitness effect of 18 009 genes tested across 323 cancer cell models. Through interactive interfaces, users can investigate data by selecting a specific gene, cancer cell model or tissue type, as well as browsing all gene fitness scores. Additionally, users can identify and rank candidate drug targets based on an established oncology target prioritization pipeline, incorporating genetic biomarkers and clinical datasets for each target, and including suitability for drug development based on pharmaceutical tractability. Data are freely available and downloadable. To enhance analyses, links to other key resources including Open Targets, COSMIC, the Cell Model Passports, UniProt and the Genomics of Drug Sensitivity in Cancer are provided. The Project Score database is a valuable new tool for investigating genetic dependencies in cancer cells and the identification of candidate oncology targets.</p>]]></description>
            <pubDate><![CDATA[2020-10-17T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The STRING database in 2021: customizable protein–protein networks, and functional characterization of user-uploaded gene/measurement sets]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609942177-e892a964-0f59-4b6d-a02d-be94c309832f/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1074</link>
            <description><![CDATA[<p class="para" id="N65541">Cellular life depends on a complex web of functional associations between biomolecules. Among these associations, protein–protein interactions are particularly important due to their versatility, specificity and adaptability. The STRING database aims to integrate all known and predicted associations between proteins, including both physical interactions as well as functional associations. To achieve this, STRING collects and scores evidence from a number of sources: (i) automated text mining of the scientific literature, (ii) databases of interaction experiments and annotated complexes/pathways, (iii) computational interaction predictions from co-expression and from conserved genomic context and (iv) systematic transfers of interaction evidence from one organism to another. STRING aims for wide coverage; the upcoming version 11.5 of the resource will contain more than 14 000 organisms. In this update paper, we describe changes to the text-mining system, a new scoring-mode for physical interactions, as well as extensive user interface features for customizing, extending and sharing protein networks. In addition, we describe how to query STRING with genome-wide, experimental data, including the automated detection of enriched functionalities and potential biases in the user's query data. The STRING resource is available online, at https://string-db.org/.</p>]]></description>
            <pubDate><![CDATA[2020-11-25T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[CoV3D: a database of high resolution coronavirus protein structures]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609929609-36764464-ac2b-4e49-a465-5eb37a91dcf3/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa731</link>
            <description><![CDATA[<p class="para" id="N65541">SARS-CoV-2, the etiologic agent of COVID-19, exemplifies the general threat to global health posed by coronaviruses. The urgent need for effective vaccines and therapies is leading to a rapid rise in the number of high resolution structures of SARS-CoV-2 proteins that collectively reveal a map of virus vulnerabilities. To assist structure-based design of vaccines and therapeutics against SARS-CoV-2 and other coronaviruses, we have developed CoV3D, a database and resource for coronavirus protein structures, which is updated on a weekly basis. CoV3D provides users with comprehensive sets of structures of coronavirus proteins and their complexes with antibodies, receptors, and small molecules. Integrated molecular viewers allow users to visualize structures of the spike glycoprotein, which is the major target of neutralizing antibodies and vaccine design efforts, as well as sets of spike-antibody complexes, spike sequence variability, and known polymorphisms. In order to aid structure-based design and analysis of the spike glycoprotein, CoV3D permits visualization and download of spike structures with modeled N-glycosylation at known glycan sites, and contains structure-based classification of spike conformations, generated by unsupervised clustering. CoV3D can serve the research community as a centralized reference and resource for spike and other coronavirus protein structures, and is available at: https://cov3d.ibbr.umd.edu.</p><p class="para" id="N65542">
<div class="section" id="ga1"><div class="img"><div class="imgeVideo"><div class="img-fullscreenIcon" onClick="javascript:showImageContent('ga1');"><img src="/public/images/journalImg/fullscreen.png"/></div><div class="imageVideo"><img src="/dataresources/secured/content-1765609929609-36764464-ac2b-4e49-a465-5eb37a91dcf3/assets/gkaa731gra1.jpg" alt="CoV3D: a database of high resolution coronavirus protein structures."/></div></div><div class="imgeVideoCaption" id="N65544"><div class="captionTitle">Graphical Abstract</div><div class="captionText">                                      CoV3D: a database of high resolution coronavirus protein structures.</div></div></div></div>
</p>]]></description>
            <pubDate><![CDATA[2020-09-05T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[OncoVar: an integrated database and analysis platform for oncogenic driver variants in cancers]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609920874-6b487f90-d94f-4ed2-8cc7-ae4e01e8f788/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1033</link>
            <description><![CDATA[<p class="para" id="N65541">The prevalence of neutral mutations in cancer cell population impedes the distinguishing of cancer-causing driver mutations from passenger mutations. To systematically prioritize the oncogenic ability of somatic mutations and cancer genes, we constructed a useful platform, OncoVar (https://oncovar.org/), which employed published bioinformatics algorithms and incorporated known driver events to identify driver mutations and driver genes. We identified 20 162 cancer driver mutations, 814 driver genes and 2360 pathogenic pathways with high-confidence by reanalyzing 10 769 exomes from 33 cancer types in The Cancer Genome Atlas (TCGA) and 1942 genomes from 18 cancer types in International Cancer Genome Consortium (ICGC). OncoVar provides four points of view, ‘Mutation’, ‘Gene’, ‘Pathway’ and ‘Cancer’, to help researchers to visualize the relationships between cancers and driver variants. Importantly, identification of actionable driver alterations provides promising druggable targets and repurposing opportunities of combinational therapies. OncoVar provides a user-friendly interface for browsing, searching and downloading somatic driver mutations, driver genes and pathogenic pathways in various cancer types. This platform will facilitate the identification of cancer drivers across individual cancer cohorts and helps to rank mutations or genes for better decision-making among clinical oncologists, cancer researchers and the broad scientific community interested in cancer precision medicine.</p>]]></description>
            <pubDate><![CDATA[2020-11-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[LncExpDB: an expression database of human long non-coding RNAs]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609866860-22570d17-7020-43c5-89c0-e50983c27e51/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa850</link>
            <description><![CDATA[<p class="para" id="N65541">Expression profiles of long non-coding RNAs (lncRNAs) across diverse biological conditions provide significant insights into their biological functions, interacting targets as well as transcriptional reliability. However, there lacks a comprehensive resource that systematically characterizes the expression landscape of human lncRNAs by integrating their expression profiles across a wide range of biological conditions. Here, we present LncExpDB (https://bigd.big.ac.cn/lncexpdb), an expression database of human lncRNAs that is devoted to providing comprehensive expression profiles of lncRNA genes, exploring their expression features and capacities, identifying featured genes with potentially important functions, and building interactions with protein-coding genes across various biological contexts/conditions. Based on comprehensive integration and stringent curation, LncExpDB currently houses expression profiles of 101 293 high-quality human lncRNA genes derived from 1977 samples of 337 biological conditions across nine biological contexts. Consequently, LncExpDB estimates lncRNA genes’ expression reliability and capacities, identifies 25 191 featured genes, and further obtains 28 443 865 lncRNA-mRNA interactions. Moreover, user-friendly web interfaces enable interactive visualization of expression profiles across various conditions and easy exploration of featured lncRNAs and their interacting partners in specific contexts. Collectively, LncExpDB features comprehensive integration and curation of lncRNA expression profiles and thus will serve as a fundamental resource for functional studies on human lncRNAs.</p>]]></description>
            <pubDate><![CDATA[2020-10-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The European Nucleotide Archive in 2020]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609825854-f1528918-0316-471e-bca5-0b5e8e3cb076/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1028</link>
            <description><![CDATA[<p class="para" id="N65541">The European Nucleotide Archive (ENA; https://www.ebi.ac.uk/ena), provided by the European Molecular Biology Laboratory's European Bioinformatics Institute (EMBL-EBI), has for almost forty years continued in its mission to freely archive and present the world's public sequencing data for the benefit of the entire scientific community and for the acceleration of the global research effort. Here we highlight the major developments to ENA services and content in 2020, focussing in particular on the recently released updated ENA browser, modernisation of our release process and our data coordination collaborations with specific research communities.</p>]]></description>
            <pubDate><![CDATA[2020-11-11T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[DIGGER: exploring the functional role of alternative splicing in protein interactions]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609820903-ac584128-37f8-4465-90d8-a715dbf6df4f/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa768</link>
            <description><![CDATA[<p class="para" id="N65541">Alternative splicing plays a major role in regulating the functional repertoire of the proteome. However, isoform-specific effects to protein-protein interactions (PPIs) are usually overlooked, making it impossible to judge the functional role of individual exons on a systems biology level. We overcome this barrier by integrating protein-protein interactions, domain-domain interactions and residue-level interactions information to lift exon expression analysis to a network level. Our user-friendly database DIGGER is available at https://exbio.wzw.tum.de/digger and allows users to seamlessly switch between isoform and exon-centric views of the interactome and to extract sub-networks of relevant isoforms, making it an essential resource for studying mechanistic consequences of alternative splicing.</p>]]></description>
            <pubDate><![CDATA[2020-09-25T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[PAGER-CoV: a comprehensive collection of pathways, annotated gene-lists and gene signatures for coronavirus disease studies]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609815586-d4a91df2-1533-4a58-8036-58773423a35b/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1094</link>
            <description><![CDATA[<p class="para" id="N65541">
<b>PAGER-CoV</b> (http://discovery.informatics.uab.edu/PAGER-CoV/) is a new web-based database that can help biomedical researchers interpret coronavirus-related functional genomic study results in the context of curated knowledge of host viral infection, inflammatory response, organ damage, and tissue repair. The new database consists of 11 835 <b>PAG</b>s (<span style="text-decoration: underline">P</span>athways, <span style="text-decoration: underline">A</span>nnotated gene-lists, or <span style="text-decoration: underline">G</span>ene signatures) from 33 public data sources. Through the web user interface, users can search by a query gene or a query term and retrieve significantly matched PAGs with all the curated information. Users can navigate from a PAG of interest to other related PAGs through either shared PAG-to-PAG co-membership relationships or PAG-to-PAG regulatory relationships, totaling 19 996 993. Users can also retrieve enriched PAGs from an input list of COVID-19 functional study result genes, customize the search data sources, and export all results for subsequent offline data analysis. In a case study, we performed a gene set enrichment analysis (GSEA) of a COVID-19 RNA-seq data set from the Gene Expression Omnibus database. Compared with the results using the standard PAGER database, PAGER-CoV allows for more sensitive matching of known immune-related gene signatures. We expect PAGER-CoV to be invaluable for biomedical researchers to find molecular biology mechanisms and tailored therapeutics to treat COVID-19 patients.</p>]]></description>
            <pubDate><![CDATA[2020-11-27T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[AcrDB: a database of anti-CRISPR operons in prokaryotes and viruses]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609793973-8973b6b0-c3b9-4a04-808c-82c88ae9d218/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa857</link>
            <description><![CDATA[<p class="para" id="N65541">CRISPR–Cas is an anti-viral mechanism of prokaryotes that has been widely adopted for genome editing. To make CRISPR–Cas genome editing more controllable and safer to use, anti-CRISPR proteins have been recently exploited to prevent excessive/prolonged Cas nuclease cleavage. Anti-CRISPR (Acr) proteins are encoded by (pro)phages/(pro)viruses, and have the ability to inhibit their host's CRISPR–Cas systems. We have built an online database AcrDB (http://bcb.unl.edu/AcrDB) by scanning ∼19 000 genomes of prokaryotes and viruses with AcrFinder, a recently developed Acr-Aca (Acr-associated regulator) operon prediction program. Proteins in Acr-Aca operons were further processed by two machine learning-based programs (AcRanker and PaCRISPR) to obtain numerical scores/ranks. Compared to other anti-CRISPR databases, AcrDB has the following unique features: (i) It is a genome-scale database with the largest collection of data (39 799 Acr-Aca operons containing Aca or Acr homologs); (ii) It offers a user-friendly web interface with various functions for browsing, graphically viewing, searching, and batch downloading Acr-Aca operons; (iii) It focuses on the genomic context of <i>Acr</i> and <i>Aca</i> candidates instead of individual Acr protein family and (iv) It collects data with three independent programs each having a unique data mining algorithm for cross validation. AcrDB will be a valuable resource to the anti-CRISPR research community.</p>]]></description>
            <pubDate><![CDATA[2020-10-17T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[CRISP-view: a database of functional genetic screens spanning multiple phenotypes]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609773320-4144750e-a400-4f08-bdbd-b6cbcc10019a/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa809</link>
            <description><![CDATA[<p class="para" id="N65541">High-throughput genetic screening based on CRISPR/Cas9 or RNA-interference (RNAi) enables the exploration of genes associated with the phenotype of interest on a large scale. The rapid accumulation of public available genetic screening data provides a wealth of knowledge about genotype-to-phenotype relationships and a valuable resource for the systematic analysis of gene functions. Here we present CRISP-view, a comprehensive database of CRISPR/Cas9 and RNAi screening datasets that span multiple phenotypes, including <i>in vitro</i> and <i>in vivo</i> cell proliferation and viability, response to cancer immunotherapy, virus response, protein expression, etc. By 22 September 2020, CRISP-view has collected 10 321 human samples and 825 mouse samples from 167 papers. All the datasets have been curated, annotated, and processed by a standard MAGeCK-VISPR analysis pipeline with quality control (QC) metrics. We also developed a user-friendly webserver to visualize, explore, and search these datasets. The webserver is freely available at http://crispview.weililab.org.</p>]]></description>
            <pubDate><![CDATA[2020-10-03T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[ADPriboDB 2.0: an updated database of ADP-ribosylated proteins]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609753653-a927145a-cdce-43fa-b81c-21171af3cb89/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa941</link>
            <description><![CDATA[<p class="para" id="N65541">ADP-ribosylation is a protein modification responsible for biological processes such as DNA repair, RNA regulation, cell cycle and biomolecular condensate formation. Dysregulation of ADP-ribosylation is implicated in cancer, neurodegeneration and viral infection. We developed ADPriboDB (adpribodb.leunglab.org) to facilitate studies in uncovering insights into the mechanisms and biological significance of ADP-ribosylation. ADPriboDB 2.0 serves as a one-stop repository comprising 48 346 entries and 9097 ADP-ribosylated proteins, of which 6708 were newly identified since the original database release. In this updated version, we provide information regarding the sites of ADP-ribosylation in 32 946 entries. The wealth of information allows us to interrogate existing databases or newly available data. For example, we found that ADP-ribosylated substrates are significantly associated with the recently identified human protein interaction networks associated with SARS-CoV-2, which encodes a conserved protein domain called macrodomain that binds and removes ADP-ribosylation. In addition, we create a new interactive tool to visualize the local context of ADP-ribosylation, such as structural and functional features as well as other post-translational modifications (e.g. phosphorylation, methylation and ubiquitination). This information provides opportunities to explore the biology of ADP-ribosylation and generate new hypotheses for experimental testing.</p>]]></description>
            <pubDate><![CDATA[2020-11-02T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Europe PMC in 2020]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609747802-c459c8b6-e220-4f71-878a-58e462c5caf3/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa994</link>
            <description><![CDATA[<p class="para" id="N65541">Europe PMC (https://europepmc.org) is a database of research articles, including peer reviewed full text articles and abstracts, and preprints - all freely available for use via website, APIs and bulk download. This article outlines new developments since 2017 where work has focussed on three key areas: (i) Europe PMC has added to its core content to include life science preprint abstracts and a special collection of full text of COVID-19-related preprints. Europe PMC is unique as an aggregator of biomedical preprints alongside peer-reviewed articles, with over 180 000 preprints available to search. (ii) Europe PMC has significantly expanded its links to content related to the publications, such as links to Unpaywall, providing wider access to full text, preprint peer-review platforms, all major curated data resources in the life sciences, and experimental protocols. The redesigned Europe PMC website features the PubMed abstract and corresponding PMC full text merged into one article page; there is more evident and user-friendly navigation within articles and to related content, plus a figure browse feature. (iii) The expanded annotations platform offers ∼1.3 billion text mined biological terms and concepts sourced from 10 providers and over 40 global data resources.</p>]]></description>
            <pubDate><![CDATA[2020-11-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The European Bioinformatics Institute: empowering cooperation in response to a global health crisis]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609738204-0bcb2889-96ab-427c-af4d-0946ed2716e8/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1077</link>
            <description><![CDATA[<p class="para" id="N65541">The European Bioinformatics Institute (EMBL-EBI; https://www.ebi.ac.uk/) provides freely available data and bioinformatics services to the scientific community, alongside its research activity and training provision. The 2020 COVID-19 pandemic has brought to the forefront a need for the scientific community to work even more cooperatively to effectively tackle a global health crisis. EMBL-EBI has been able to build on its position to contribute to the fight against COVID-19 in a number of ways. Firstly, EMBL-EBI has used its infrastructure, expertise and network of international collaborations to help build the European COVID-19 Data Platform (https://www.covid19dataportal.org/), which brings together COVID-19 biomolecular data and connects it to researchers, clinicians and public health professionals. By September 2020, the COVID-19 Data Platform has integrated in excess of 170 000 COVID-19 biomolecular data and literature records, collected through a number of EMBL-EBI resources. Secondly, EMBL-EBI has strived to continue its support of the life science communities through the crisis, with updated Training provision and improved service provision throughout its resources. The COVID-19 pandemic has highlighted the importance of EMBL-EBI’s core principles, including international cooperation, resource sharing and central data brokering, and has further empowered scientific cooperation.</p>]]></description>
            <pubDate><![CDATA[2020-11-27T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[tRFtarget: a database for transfer RNA-derived fragment targets]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609717546-a952ee47-160c-4a58-9a45-2a5b3c4ae386/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa831</link>
            <description><![CDATA[<p class="para" id="N65541">Transfer RNA-derived fragments (tRFs) are a new class of small non-coding RNAs and play important roles in biological and physiological processes. Prediction of tRF target genes and binding sites is crucial in understanding the biological functions of tRFs in the molecular mechanisms of human diseases. We developed a publicly accessible web-based database, tRFtarget (http://trftarget.net), for tRF target prediction. It contains the computationally predicted interactions between tRFs and mRNA transcripts using the two state-of-the-art prediction tools RNAhybrid and IntaRNA, including location of the binding sites on the target, the binding region, and free energy of the binding stability with graphic illustration. tRFtarget covers 936 tRFs and 135 thousand predicted targets in eight species. It allows researchers to search either target genes by tRF IDs or tRFs by gene symbols/transcript names. We also integrated the manually curated experimental evidence of the predicted interactions into the database. Furthermore, we provided a convenient link to the DAVID<sup>®</sup> web server to perform downstream functional pathway analysis and gene ontology annotation on the predicted target genes. This database provides useful information for the scientific community to experimentally validate tRF target genes and facilitate the investigation of the molecular functions and mechanisms of tRFs.</p>]]></description>
            <pubDate><![CDATA[2020-10-09T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[DBAASP v3: database of antimicrobial/cytotoxic activity and structure of peptides as a resource for development of new therapeutics]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609703446-79bebc4b-0b58-4eda-a915-cdd25f2e60d0/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa991</link>
            <description><![CDATA[<p class="para" id="N65541">The Database of Antimicrobial Activity and Structure of Peptides (DBAASP) is an open-access, comprehensive database containing information on amino acid sequences, chemical modifications, 3D structures, bioactivities and toxicities of peptides that possess antimicrobial properties. DBAASP is updated continuously, and at present, version 3.0 (DBAASP v3) contains &gt;15 700 entries (8000 more than the previous version), including &gt;14 500 monomers and nearly 400 homo- and hetero-multimers. Of the monomeric antimicrobial peptides (AMPs), &gt;12 000 are synthetic, about 2700 are ribosomally synthesized, and about 170 are non-ribosomally synthesized. Approximately 3/4 of the entries were added after the initial release of the database in 2014 reflecting the recent sharp increase in interest in AMPs. Despite the increased interest, adoption of peptide antimicrobials in clinical practice is still limited as a consequence of several factors including side effects, problems with bioavailability and high production costs. To assist in developing and optimizing <i>de novo</i> peptides with desired biological activities, DBAASP offers several tools including a sophisticated multifactor analysis of relevant physicochemical properties. Furthermore, DBAASP has implemented a structure modelling pipeline that automates the setup, execution and upload of molecular dynamics (MD) simulations of database peptides. At present, &gt;3200 peptides have been populated with MD trajectories and related analyses that are both viewable within the web browser and available for download. More than 400 DBAASP entries also have links to experimentally determined structures in the Protein Data Bank. DBAASP v3 is freely accessible at http://dbaasp.org.</p>]]></description>
            <pubDate><![CDATA[2020-11-05T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[CSEA-DB: an omnibus for human complex trait and cell type associations]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609692071-52d4ea36-0f02-4990-aedb-feea5439dd84/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1064</link>
            <description><![CDATA[<p class="para" id="N65541">During the past decade, genome-wide association studies (GWAS) have identified many genetic variants with susceptibility to several thousands of complex diseases or traits. The genetic regulation of gene expression is highly tissue-specific and cell type-specific. Recently, single-cell technology has paved the way to dissect cellular heterogeneity in human tissues. Here, we present a reference database for GWAS trait-associated cell type-specificity, named Cell type-Specific Enrichment Analysis DataBase (CSEA-DB, available at https://bioinfo.uth.edu/CSEADB/). Specifically, we curated total of 5120 GWAS summary statistics data for a wide range of human traits and diseases followed by rigorous quality control. We further collected &gt;900 000 cells from the leading consortia such as Human Cell Landscape, Human Cell Atlas, and extensive literature mining, including 752 tissue cell types from 71 adult and fetal tissues across 11 human organ systems. The tissues and cell types were annotated with Uberon and Cell Ontology. By applying our deTS algorithm, we conducted 10 250 480 times of trait-cell type associations, reporting a total of 598 (11.68%) GWAS traits with at least one significantly associated cell type. In summary, CSEA-DB could serve as a repository of association map for human complex traits and their underlying cell types, manually curated GWAS, and single-cell transcriptome resources.</p>]]></description>
            <pubDate><![CDATA[2020-11-19T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[BastionHub: a universal platform for integrating and analyzing substrates secreted by Gram-negative bacteria]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609649825-f8f290b8-e455-49fb-b925-d0cb886d998c/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa899</link>
            <description><![CDATA[<p class="para" id="N65541">Gram-negative bacteria utilize secretion systems to export substrates into their surrounding environment or directly into neighboring cells. These substrates are proteins that function to promote bacterial survival: by facilitating nutrient collection, disabling competitor species or, for pathogens, to disable host defenses. Following a rapid development of computational techniques, a growing number of substrates have been discovered and subsequently validated by wet lab experiments. To date, several online databases have been developed to catalogue these substrates but they have limited user options for in-depth analysis, and typically focus on a single type of secreted substrate. We therefore developed a universal platform, BastionHub, that incorporates extensive functional modules to facilitate substrate analysis and integrates the five major Gram-negative secreted substrate types (i.e. from types I–IV and VI secretion systems). To our knowledge, BastionHub is not only the most comprehensive online database available, it is also the first to incorporate substrates secreted by type I or type II secretion systems. By providing the most up-to-date details of secreted substrates and state-of-the-art prediction and visualized relationship analysis tools, BastionHub will be an important platform that can assist biologists in uncovering novel substrates and formulating new hypotheses. BastionHub is freely available at http://bastionhub.erc.monash.edu/.</p><p class="para" id="N65542">
<div class="section" id="ga1"><div class="img"><div class="imgeVideo"><div class="img-fullscreenIcon" onClick="javascript:showImageContent('ga1');"><img src="/public/images/journalImg/fullscreen.png"/></div><div class="imageVideo"><img src="/dataresources/secured/content-1765609649825-f8f290b8-e455-49fb-b925-d0cb886d998c/assets/gkaa899gra1.jpg" alt="BastionHub: a universal platform for integrating and analyzing substrates secreted by Gram-negative bacteria."/></div></div><div class="imgeVideoCaption" id="N65544"><div class="captionTitle">Graphical Abstract</div><div class="captionText">                                      BastionHub: a universal platform for integrating and analyzing substrates secreted by Gram-negative bacteria.</div></div></div></div>
</p>]]></description>
            <pubDate><![CDATA[2020-10-21T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[DNAmoreDB, a database of DNAzymes]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609620807-87c2b368-b8e1-4a5a-9bf6-60105fa7b80d/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa867</link>
            <description><![CDATA[<p class="para" id="N65541">Deoxyribozymes, DNA enzymes or simply DNAzymes are single-stranded oligo-deoxyribonucleotide molecules that, like proteins and ribozymes, possess the ability to perform catalysis. Although DNAzymes have not yet been found in living organisms, they have been isolated in the laboratory through in vitro selection. The selected DNAzyme sequences have the ability to catalyze a broad range of chemical reactions, utilizing DNA, RNA, peptides or small organic compounds as substrates. DNAmoreDB is a comprehensive database resource for DNAzymes that collects and organizes the following types of information: sequences, conditions of the selection procedure, catalyzed reactions, kinetic parameters, substrates, cofactors, structural information whenever available, and literature references. Currently, DNAmoreDB contains information about DNAzymes that catalyze 20 different reactions. We included a submission form for new data, a REST-based API system that allows users to retrieve the database contents in a machine-readable format, and keyword and BLASTN search features. The database is publicly available at https://www.genesilico.pl/DNAmoreDB/.</p><p class="para" id="N65542">
<div class="section" id="ga1"><div class="img"><div class="imgeVideo"><div class="img-fullscreenIcon" onClick="javascript:showImageContent('ga1');"><img src="/public/images/journalImg/fullscreen.png"/></div><div class="imageVideo"><img src="/dataresources/secured/content-1765609620807-87c2b368-b8e1-4a5a-9bf6-60105fa7b80d/assets/gkaa867gra1.jpg" alt="DNAmoreDB, a comprehensive database resource for DNAzymes."/></div></div><div class="imgeVideoCaption" id="N65544"><div class="captionTitle">Graphical Abstract</div><div class="captionText">                                      DNAmoreDB, a comprehensive database resource for DNAzymes.</div></div></div></div>
</p>]]></description>
            <pubDate><![CDATA[2020-10-14T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[GESS: a database of global evaluation of SARS-CoV-2/hCoV-19 sequences]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609616246-bece6de7-238e-4bf4-ba9b-0e67cb663173/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa808</link>
            <description><![CDATA[<p class="para" id="N65541">The COVID-19 outbreak has become a global emergency since December 2019. Analysis of SARS-CoV-2 sequences can uncover single nucleotide variants (SNVs) and corresponding evolution patterns. The <span style="text-decoration: underline">G</span>lobal <span style="text-decoration: underline">E</span>valuation of <span style="text-decoration: underline">S</span>ARS-CoV-2/hCoV-19 <span style="text-decoration: underline">S</span>equences (GESS, https://wan-bioinfo.shinyapps.io/GESS/) is a resource to provide comprehensive analysis results based on tens of thousands of high-coverage and high-quality SARS-CoV-2 complete genomes. The database allows user to browse, search and download SNVs at any individual or multiple SARS-CoV-2 genomic positions, or within a chosen genomic region or protein, or in certain country/area of interest. GESS reveals geographical distributions of SNVs around the world and across the states of USA, while exhibiting time-dependent patterns for SNV occurrences which reflect development of SARS-CoV-2 genomes. For each month, the top 100 SNVs that were firstly identified world-widely can be retrieved. GESS also explores SNVs occurring simultaneously with specific SNVs of user's interests. Furthermore, the database can be of great help to calibrate mutation rates and identify conserved genome regions. Taken together, GESS is a powerful resource and tool to monitor SARS-CoV-2 migration and evolution according to featured genomic variations. It provides potential directive information for prevalence prediction, related public health policy making, and vaccine designs.</p>]]></description>
            <pubDate><![CDATA[2020-10-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[CovalentInDB: a comprehensive database facilitating the discovery of covalent inhibitors]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609586510-e90bf525-5d2b-4ad8-9137-448ab805470b/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa876</link>
            <description><![CDATA[<p class="para" id="N65541">Inhibitors that form covalent bonds with their targets have traditionally been considered highly adventurous due to their potential off-target effects and toxicity concerns. However, with the clinical validation and approval of many covalent inhibitors during the past decade, design and discovery of novel covalent inhibitors have attracted increasing attention. A large amount of scattered experimental data for covalent inhibitors have been reported, but a resource by integrating the experimental information for covalent inhibitor discovery is still lacking. In this study, we presented Covalent Inhibitor Database (CovalentInDB), the largest online database that provides the structural information and experimental data for covalent inhibitors. CovalentInDB contains 4511 covalent inhibitors (including 68 approved drugs) with 57 different reactive warheads for 280 protein targets. The crystal structures of some of the proteins bound with a covalent inhibitor are provided to visualize the protein–ligand interactions around the binding site. Each covalent inhibitor is annotated with the structure, warhead, experimental bioactivity, physicochemical properties, etc. Moreover, CovalentInDB provides the covalent reaction mechanism and the corresponding experimental verification methods for each inhibitor towards its target. High-quality datasets are downloadable for users to evaluate and develop computational methods for covalent drug design. CovalentInDB is freely accessible at http://cadd.zju.edu.cn/cidb/.</p><p class="para" id="N65542">
<div class="section" id="ga1"><div class="img"><div class="imgeVideo"><div class="img-fullscreenIcon" onClick="javascript:showImageContent('ga1');"><img src="/public/images/journalImg/fullscreen.png"/></div><div class="imageVideo"><img src="/dataresources/secured/content-1765609586510-e90bf525-5d2b-4ad8-9137-448ab805470b/assets/gkaa876gra1.jpg" alt="The information for covalent inhibitors and related targets is deposited in the CovalentInDB database for users to browse, retrieve and download."/></div></div><div class="imgeVideoCaption" id="N65544"><div class="captionTitle">Graphical Abstract</div><div class="captionText">                                      The information for covalent inhibitors and related targets is deposited in the CovalentInDB database for users to browse, retrieve and download.</div></div></div></div>
</p>]]></description>
            <pubDate><![CDATA[2020-10-17T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[LnCeCell: a comprehensive database of predicted lncRNA-associated ceRNA networks at single-cell resolution]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609581751-5b18de8d-cc2e-49e6-afa7-520228ec533a/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1017</link>
            <description><![CDATA[<p class="para" id="N65541">Within the tumour microenvironment, cells exhibit different behaviours driven by fine-tuning of gene regulation. Identification of cellular-specific gene regulatory networks will deepen the understanding of disease pathology at single-cell resolution and contribute to the development of precision medicine. Here, we describe a database, LnCeCell (http://www.bio-bigdata.net/LnCeCell/ or http://bio-bigdata.hrbmu.edu.cn/LnCeCell/), which aims to document cellular-specific long non-coding RNA (lncRNA)-associated competing endogenous RNA (ceRNA) networks for personalised characterisation of diseases based on the ‘One Cell, One World’ theory. LnCeCell is curated with cellular-specific ceRNA regulations from &gt;94 000 cells across 25 types of cancers and provides &gt;9000 experimentally supported lncRNA biomarkers, associated with tumour metastasis, recurrence, prognosis, circulation, drug resistance, etc. For each cell, LnCeCell illustrates a global map of ceRNA sub-cellular locations, which have been manually curated from the literature and related data sources, and portrays a functional state atlas for a single cancer cell. LnCeCell also provides several flexible tools to infer ceRNA functions based on a specific cellular background. LnCeCell serves as an important resource for investigating the gene regulatory networks within a single cell and can help researchers understand the regulatory mechanisms underlying complex microbial ecosystems and individual phenotypes.</p>]]></description>
            <pubDate><![CDATA[2020-11-21T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[MarkerDB: an online database of molecular biomarkers]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609559654-36374233-ea6c-4e39-8296-6d1002dc517b/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1067</link>
            <description><![CDATA[<p class="para" id="N65541">MarkerDB is a freely available electronic database that attempts to consolidate information on all known clinical and a selected set of pre-clinical molecular biomarkers into a single resource. The database includes four major types of molecular biomarkers (chemical, protein, DNA [genetic] and karyotypic) and four biomarker categories (diagnostic, predictive, prognostic and exposure). MarkerDB provides information such as: biomarker names and synonyms, associated conditions or pathologies, detailed disease descriptions, detailed biomarker descriptions, biomarker specificity, sensitivity and ROC curves, standard reference values (for protein and chemical markers), variants (for SNP or genetic markers), sequence information (for genetic and protein markers), molecular structures (for protein and chemical markers), tissue or biofluid sources (for protein and chemical markers), chromosomal location and structure (for genetic and karyotype markers), clinical approval status and relevant literature references. Users can browse the data by conditions, condition categories, biomarker types, biomarker categories or search by sequence similarity through the advanced search function. Currently, the database contains 142 protein biomarkers, 1089 chemical biomarkers, 154 karyotype biomarkers and 26 374 genetic markers. These are categorized into 25 560 diagnostic biomarkers, 102 prognostic biomarkers, 265 exposure biomarkers and 6746 predictive biomarkers or biomarker panels. Collectively, these markers can be used to detect, monitor or predict 670 specific human conditions which are grouped into 27 broad condition categories. MarkerDB is available at https://markerdb.ca.</p>]]></description>
            <pubDate><![CDATA[2020-11-27T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[ProThermDB: thermodynamic database for proteins and mutants revisited after 15 years]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609555679-362ba0a0-eea9-4438-af47-b800f0d1d4c3/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1035</link>
            <description><![CDATA[<p class="para" id="N65541">ProThermDB is an updated version of the thermodynamic database for proteins and mutants (ProTherm), which has ∼31 500 data on protein stability, an increase of 84% from the previous version. It contains several thermodynamic parameters such as melting temperature, free energy obtained with thermal and denaturant denaturation, enthalpy change and heat capacity change along with experimental methods and conditions, sequence, structure and literature information. Besides, the current version of the database includes about 120 000 thermodynamic data obtained for different organisms and cell lines, which are determined by recent high throughput proteomics techniques using whole-cell approaches. In addition, we provided a graphical interface for visualization of mutations at sequence and structure levels. ProThermDB is cross-linked with other relevant databases, PDB, UniProt, PubMed etc. It is freely available at https://web.iitm.ac.in/bioinfo2/prothermdb/index.html without any login requirements. It is implemented in Python, HTML and JavaScript, and supports the latest versions of major browsers, such as Firefox, Chrome and Safari.</p>]]></description>
            <pubDate><![CDATA[2020-11-16T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[LectomeXplore, an update of UniLectin for the discovery of carbohydrate-binding proteins based on a new lectin classification]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609525898-d1a29522-89bd-4290-83ee-7bfb9fb9c782/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1019</link>
            <description><![CDATA[<p class="para" id="N65541">Lectins are non-covalent glycan-binding proteins mediating cellular interactions but their annotation in newly sequenced organisms is lacking. The limited size of functional domains and the low level of sequence similarity challenge usual bioinformatics tools. The identification of lectin domains in proteomes requires the manual curation of sequence alignments based on structural folds. A new lectin classification is proposed. It is built on three levels: (i) 35 lectin domain folds, (ii) 109 classes of lectins sharing at least 20% sequence similarity and (iii) 350 families of lectins sharing at least 70% sequence similarity. This information is compiled in the UniLectin platform that includes the previously described UniLectin3D database of curated lectin 3D structures. Since its first release, UniLectin3D has been updated with 485 additional 3D structures. The database is now complemented by two additional modules: PropLec containing predicted β-propeller lectins and LectomeXplore including predicted lectins from sequences of the NBCI-nr and UniProt for every curated lectin class. UniLectin is accessible at https://www.unilectin.eu/</p>]]></description>
            <pubDate><![CDATA[2020-11-11T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The Dark Kinase Knowledgebase: an online compendium of knowledge and experimental results of understudied kinases]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609514600-732c4859-9d46-403e-8758-7ddde2998efa/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa853</link>
            <description><![CDATA[<p class="para" id="N65541">Kinases form the backbone of numerous cell signaling pathways, with their dysfunction similarly implicated in multiple pathologies. Further facilitated by their druggability, kinases are a major focus of therapeutic development efforts in diseases such as cancer, infectious disease and autoimmune disorders. While their importance is clear, the role or biological function of nearly one-third of kinases is largely unknown. Here, we describe a data resource, the Dark Kinase Knowledgebase (DKK; https://darkkinome.org), that is specifically focused on providing data and reagents for these understudied kinases to the broader research community. Supported through NIH’s Illuminating the Druggable Genome (IDG) Program, the DKK is focused on data and knowledge generation for 162 poorly studied or ‘dark’ kinases. Types of data provided through the DKK include parallel reaction monitoring (PRM) peptides for quantitative proteomics, protein interactions, NanoBRET reagents, and kinase-specific compounds. Higher-level data is similarly being generated and consolidated such as tissue gene expression profiles and, longer-term, functional relationships derived through perturbation studies. Associated web tools that help investigators interrogate both internal and external data are also provided through the site. As an evolving resource, the DKK seeks to continually support and enhance knowledge on these potentially high-impact druggable targets.</p>]]></description>
            <pubDate><![CDATA[2020-10-20T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[GlycoPOST realizes FAIR principles for glycomics mass spectrometry data]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609511264-c3d1538f-e2b5-418a-93df-6f6df7008ab8/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1012</link>
            <description><![CDATA[<p class="para" id="N65541">For the reproducibility and sustainability of scientific research, FAIRness (Findable, Accessible, Interoperable and Re-usable), with respect to the release of raw data obtained by researchers, is one of the most important principles underpinning the future of open science. In genomics and transcriptomics, the sharing of raw data from next-generation sequencers is made possible through public repositories. In addition, in proteomics, the deposition of raw data from mass spectrometry (MS) experiments into repositories is becoming standardized. However, a standard repository for such MS data had not yet been established in glycomics. With the increasing number of glycomics MS data, therefore, we have developed GlycoPOST (https://glycopost.glycosmos.org/), a repository for raw MS data generated from glycomics experiments. In just the first year since the release of GlycoPOST, 73 projects have already been registered by researchers around the world, and the number of registered projects is continuously growing, making a significant contribution to the future FAIRness of the glycomics field. GlycoPOST is a free resource to the community and accepts (and will continue to accept in the future) raw data regardless of vendor-specific formats.</p>]]></description>
            <pubDate><![CDATA[2020-11-11T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[STAB: a spatio-temporal cell atlas of the human brain]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609482755-9cddf986-7720-4c62-8b75-86d5026dae5f/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa762</link>
            <description><![CDATA[<p class="para" id="N65541">The human brain is the most complex organ consisting of billions of neuronal and non-neuronal cells that are organized into distinct anatomical and functional regions. Elucidating the cellular and transcriptome architecture underlying the brain is crucial for understanding brain functions and brain disorders. Thanks to the single-cell RNA sequencing technologies, it is becoming possible to dissect the cellular compositions of the brain. Although great effort has been made to explore the transcriptome architecture of the human brain, a comprehensive database with dynamic cellular compositions and molecular characteristics of the human brain during the lifespan is still not available. Here, we present STAB (a <span style="text-decoration: underline">S</span>patio-<span style="text-decoration: underline">T</span>emporal cell <span style="text-decoration: underline">A</span>tlas of the human <span style="text-decoration: underline">B</span>rain), a database consists of single-cell transcriptomes across multiple brain regions and developmental periods. Right now, STAB contains single-cell gene expression profiling of 42 cell subtypes across 20 brain regions and 11 developmental periods. With STAB, the landscape of cell types and their regional heterogeneity and temporal dynamics across the human brain can be clearly seen, which can help to understand both the development of the normal human brain and the etiology of neuropsychiatric disorders. STAB is available at http://stab.comp-sysbio.org.</p>]]></description>
            <pubDate><![CDATA[2020-09-25T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The international nucleotide sequence database collaboration]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609464051-de84e24d-311d-41e4-bdf7-343157de8490/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa967</link>
            <description><![CDATA[<p class="para" id="N65541">The International Nucleotide Sequence Database Collaboration (INSDC; http://www.insdc.org/) has been the core infrastructure for collecting and providing nucleotide sequence data and metadata for &gt;30 years. Three partner organizations, the DNA Data Bank of Japan (DDBJ) at the National Institute of Genetics in Mishima, Japan; the European Nucleotide Archive (ENA) at the European Molecular Biology Laboratory's European Bioinformatics Institute (EMBL-EBI) in Hinxton, UK; and GenBank at National Center for Biotechnology Information (NCBI), National Library of Medicine, National Institutes of Health in Bethesda, Maryland, USA have been collaboratively maintaining the INSDC for the benefit of not only science but all types of community worldwide.</p>]]></description>
            <pubDate><![CDATA[2020-11-09T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[LegumeIP V3: from models to crops—an integrative gene discovery platform for translational genomics in legumes]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609460107-754c7580-6a40-4a05-b8a5-48640d44c22c/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa976</link>
            <description><![CDATA[<p class="para" id="N65541">Legumes have contributed to human health, sustainable food and feed production worldwide for centuries. The study of model legumes has played vital roles in deciphering key genes, pathways, and networks regulating biological mechanisms and agronomic traits. Along with emerging breeding technology such as genome editing, translation of the knowledge gained from model plants to crops is in high demand. The updated database (V3) was redesigned for translational genomics targeting the discovery of novel key genes in less-studied non-model legume crops by referring to the knowledge gained in model legumes. The database contains genomic data for all 22 included species, and transcriptomic data covering thousands of RNA-seq samples mostly from model species. The rich biological data and analytic tools for gene expression and pathway analyses can be used to decipher critical genes, pathways, and networks in model legumes. The integrated comparative genomic functions further facilitate the translation of this knowledge to legume crops. Therefore, the database will be a valuable resource to identify important genes regulating specific biological mechanisms or agronomic traits in the non-model yet economically significant legume crops. LegumeIP V3 is available free to the public at https://plantgrn.noble.org/LegumeIP. Access to the database does not require login, registration, or password.</p>]]></description>
            <pubDate><![CDATA[2020-11-09T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[PED in 2021: a major update of the protein ensemble database for intrinsically disordered proteins]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609447210-7f1b2195-8932-411c-a6a6-2516fcc873aa/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1021</link>
            <description><![CDATA[<p class="para" id="N65541">The Protein Ensemble Database (PED) (https://proteinensemble.org), which holds structural ensembles of intrinsically disordered proteins (IDPs), has been significantly updated and upgraded since its last release in 2016. The new version, PED 4.0, has been completely redesigned and reimplemented with cutting-edge technology and now holds about six times more data (162 versus 24 entries and 242 versus 60 structural ensembles) and a broader representation of state of the art ensemble generation methods than the previous version. The database has a completely renewed graphical interface with an interactive feature viewer for region-based annotations, and provides a series of descriptors of the qualitative and quantitative properties of the ensembles. High quality of the data is guaranteed by a new submission process, which combines both automatic and manual evaluation steps. A team of biocurators integrate structured metadata describing the ensemble generation methodology, experimental constraints and conditions. A new search engine allows the user to build advanced queries and search all entry fields including cross-references to IDP-related resources such as DisProt, MobiDB, BMRB and SASBDB. We expect that the renewed PED will be useful for researchers interested in the atomic-level understanding of IDP function, and promote the rational, structure-based design of IDP-targeting drugs.</p>]]></description>
            <pubDate><![CDATA[2020-12-10T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[SMART: recent updates, new developments and status in 2020]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609444292-c3e5719d-d425-4c7a-b53e-ea532168e962/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa937</link>
            <description><![CDATA[<p class="para" id="N65541">SMART (Simple Modular Architecture Research Tool) is a web resource (https://smart.embl.de) for the identification and annotation of protein domains and the analysis of protein domain architectures. SMART version 9 contains manually curated models for more than 1300 protein domains, with a topical set of 68 new models added since our last update article (<a href="#B1">1</a>). All the new models are for diverse recombinase families and subfamilies and as a set they provide a comprehensive overview of mobile element recombinases namely transposase, integrase, relaxase, resolvase, cas1 casposase and Xer like cellular recombinase. Further updates include the synchronization of the underlying protein databases with UniProt (<a href="#B2">2</a>), Ensembl (<a href="#B3">3</a>) and STRING (<a href="#B4">4</a>), greatly increasing the total number of annotated domains and other protein features available in architecture analysis mode. Furthermore, SMART’s vector-based protein display engine has been extended and updated to use the latest web technologies and the domain architecture analysis components have been optimized to handle the increased number of protein features available.</p>]]></description>
            <pubDate><![CDATA[2020-10-26T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[BacWGSTdb 2.0: a one-stop repository for bacterial whole-genome sequence typing and source tracking]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609420633-5261945c-f188-4386-b660-6ad95afe899d/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa821</link>
            <description><![CDATA[<p class="para" id="N65541">An increasing prevalence of hospital acquired infections and foodborne illnesses caused by pathogenic and multidrug-resistant bacteria has stimulated a pressing need for benchtop computational techniques to rapidly and accurately classify bacteria from genomic sequence data, and based on that, to trace the source of infection. BacWGSTdb (http://bacdb.org/BacWGSTdb) is a free publicly accessible database we have developed for bacterial whole-genome sequence typing and source tracking. This database incorporates extensive resources for bacterial genome sequencing data and the corresponding metadata, combined with specialized bioinformatics tools that enable the systematic characterization of the bacterial isolates recovered from infections. Here, we present BacWGSTdb 2.0, which encompasses several major updates, including (i) the integration of the core genome multi-locus sequence typing (cgMLST) approach, which is highly scalable and appropriate for typing isolates belonging to different lineages; (ii) the addition of a multiple genome analysis module that can process dozens of user uploaded sequences in a batch mode; (iii) a new source tracking module for comparing user uploaded plasmid sequences to those deposited in the public databases; (iv) the number of species encompassed in BacWGSTdb 2.0 has increased from 9 to 20, which represents bacterial pathogens of medical importance; (v) a newly designed, user-friendly interface and a set of visualization tools for providing a convenient platform for users are also included. Overall, the updated BacWGSTdb 2.0 bears great utility in continuing to provide users, including epidemiologists, clinicians and bench scientists, with a one-stop solution to bacterial genome sequence analysis.</p>]]></description>
            <pubDate><![CDATA[2020-10-03T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The mouse Gene Expression Database (GXD): 2021 update]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609416055-dacf7621-4260-412f-a90f-cd0ad9e2d85f/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa914</link>
            <description><![CDATA[<p class="para" id="N65541">The Gene Expression Database (GXD; www.informatics.jax.org/expression.shtml) is an extensive and well-curated community resource of mouse developmental gene expression information. For many years, GXD has collected and integrated data from RNA <i>in situ</i> hybridization, immunohistochemistry, RT-PCR, northern blot, and western blot experiments through curation of the scientific literature and by collaborations with large-scale expression projects. Since our last report in 2019, we have continued to acquire these classical types of expression data; developed a searchable index of RNA-Seq and microarray experiments that allows users to quickly and reliably find specific mouse expression studies in ArrayExpress (https://www.ebi.ac.uk/arrayexpress/) and GEO (https://www.ncbi.nlm.nih.gov/geo/); and expanded GXD to include RNA-Seq data. Uniformly processed RNA-Seq data are imported from the EBI Expression Atlas and then integrated with the other types of expression data in GXD, and with the genetic, functional, phenotypic and disease-related information in Mouse Genome Informatics (MGI). This integration has made the RNA-Seq data accessible via GXD’s enhanced searching and filtering capabilities. Further, we have embedded the Morpheus heat map utility into the GXD user interface to provide additional tools for display and analysis of RNA-Seq data, including heat map visualization, sorting, filtering, hierarchical clustering, nearest neighbors analysis and visual enrichment.</p>]]></description>
            <pubDate><![CDATA[2020-10-26T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Planet Microbe: a platform for marine microbiology to discover and analyze interconnected ‘omics and environmental data]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609314975-12e92a20-a482-45b8-b927-8b288ce651b0/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa637</link>
            <description><![CDATA[<p class="para" id="N65541">In recent years, large-scale oceanic sequencing efforts have provided a deeper understanding of marine microbial communities and their dynamics. These research endeavors require the acquisition of complex and varied datasets through large, interdisciplinary and collaborative efforts. However, no unifying framework currently exists for the marine science community to integrate sequencing data with physical, geological, and geochemical datasets. Planet Microbe is a web-based platform that enables data discovery from curated historical and on-going oceanographic sequencing efforts. In Planet Microbe, each ‘omics sample is linked with other biological and physiochemical measurements collected for the same water samples or during the same sample collection event, to provide a broader environmental context. This work highlights the need for curated aggregation efforts that can enable new insights into high-quality metagenomic datasets. Planet Microbe is freely accessible from https://www.planetmicrobe.org/.</p>]]></description>
            <pubDate><![CDATA[2020-07-31T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[TCRdb: a comprehensive database for T-cell receptor sequences with powerful search function]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609289041-3b7547f0-8a42-4de8-b59e-ef81fb428102/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa796</link>
            <description><![CDATA[<p class="para" id="N65541">T cells and the T-cell receptor (TCR) repertoire play pivotal roles in immune response and immunotherapy. TCR sequencing (TCR-Seq) technology has enabled accurate profiling TCR repertoire and currently a large number of TCR-Seq data are available in public. Based on the urgent need to effectively re-use these data, we developed TCRdb, a comprehensive human TCR sequences database, by a uniform pipeline to characterize TCR sequences on TCR-Seq data. TCRdb contains more than 277 million highly reliable TCR sequences from over 8265 TCR-Seq samples across hundreds of tissues/clinical conditions/cell types. The unique features of TCRdb include: (i) comprehensive and reliable sequences for TCR repertoire in different samples generated by a strict and uniform pipeline of TCRdb; (ii) powerful search function, allowing users to identify their interested TCR sequences in different conditions; (iii) categorized sample metadata, enabling comparison of TCRs in different sample types; (iv) interactive data visualization charts, describing the TCR repertoire in TCR diversity, length distribution and V-J gene utilization. The TCRdb database is freely available at http://bioinfo.life.hust.edu.cn/TCRdb/ and will be a useful resource in the research and application community of T cell immunology.</p>]]></description>
            <pubDate><![CDATA[2020-09-29T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[RMDisease: a database of genetic variants that affect RNA modifications, with implications for epitranscriptome pathogenesis]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609272507-482c7f88-ac32-45d7-ab8a-bee2dc305426/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa790</link>
            <description><![CDATA[<p class="para" id="N65541">Deciphering the biological impacts of millions of single nucleotide variants remains a major challenge. Recent studies suggest that RNA modifications play versatile roles in essential biological mechanisms, and are closely related to the progression of various diseases including multiple cancers. To comprehensively unveil the association between disease-associated variants and their epitranscriptome disturbance, we built RMDisease, a database of genetic variants that can affect RNA modifications. By integrating the prediction results of 18 different RNA modification prediction tools and also 303,426 experimentally-validated RNA modification sites, RMDisease identified a total of 202,307 human SNPs that may affect (add or remove) sites of eight types of RNA modifications (m<sup>6</sup>A, m<sup>5</sup>C, m<sup>1</sup>A, m<sup>5</sup>U, Ψ, m<sup>6</sup>Am, m<sup>7</sup>G and Nm). These include 4,289 disease-associated variants that may imply disease pathogenesis functioning at the epitranscriptome layer. These SNPs were further annotated with essential information such as post-transcriptional regulations (sites for miRNA binding, interaction with RNA-binding proteins and alternative splicing) revealing putative regulatory circuits. A convenient graphical user interface was constructed to support the query, exploration and download of the relevant information. RMDisease should make a useful resource for studying the epitranscriptome impact of genetic variants via multiple RNA modifications with emphasis on their potential disease relevance. RMDisease is freely accessible at: www.xjtlu.edu.cn/biologicalsciences/rmd.</p>]]></description>
            <pubDate><![CDATA[2020-10-03T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[TISCH: a comprehensive web resource enabling interactive single-cell transcriptome visualization of tumor microenvironment]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609257516-0be66754-6a32-422a-84e0-ab7c5ce5e980/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1020</link>
            <description><![CDATA[<p class="para" id="N65541">Cancer immunotherapy targeting co-inhibitory pathways by checkpoint blockade shows remarkable efficacy in a variety of cancer types. However, only a minority of patients respond to treatment due to the stochastic heterogeneity of tumor microenvironment (TME). Recent advances in single-cell RNA-seq technologies enabled comprehensive characterization of the immune system heterogeneity in tumors but posed computational challenges on integrating and utilizing the massive published datasets to inform immunotherapy. Here, we present Tumor Immune Single Cell Hub (TISCH, http://tisch.comp-genomics.org), a large-scale curated database that integrates single-cell transcriptomic profiles of nearly 2 million cells from 76 high-quality tumor datasets across 27 cancer types. All the data were uniformly processed with a standardized workflow, including quality control, batch effect removal, clustering, cell-type annotation, malignant cell classification, differential expression analysis and functional enrichment analysis. TISCH provides interactive gene expression visualization across multiple datasets at the single-cell level or cluster level, allowing systematic comparison between different cell-types, patients, tissue origins, treatment and response groups, and even different cancer-types. In summary, TISCH provides a user-friendly interface for systematically visualizing, searching and downloading gene expression atlas in the TME from multiple cancer types, enabling fast, flexible and comprehensive exploration of the TME.</p>]]></description>
            <pubDate><![CDATA[2020-11-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The Transporter Classification Database (TCDB): 2021 update]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609184305-118692e3-7195-4c66-b879-ba4e16fb3948/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1004</link>
            <description><![CDATA[<p class="para" id="N65541">The Transporter Classification Database (TCDB; tcdb.org) is a freely accessible reference resource, which provides functional, structural, mechanistic, medical and biotechnological information about transporters from organisms of all types. TCDB is the only transport protein classification database adopted by the International Union of Biochemistry and Molecular Biology (IUBMB) and now (October 1, 2020) consists of 20 653 proteins classified in 15 528 non-redundant transport systems with 1567 tabulated 3D structures, 18 336 reference citations describing 1536 transporter families, of which 26% are members of 82 recognized superfamilies. Overall, this is an increase of over 50% since the last published update of the database in 2016. This comprehensive update of the database contents and features include (i) adoption of a chemical ontology for substrates of transporters, (ii) inclusion of new superfamilies, (iii) a domain-based characterization of transporter families for the identification of new members as well as functional and evolutionary relationships between families, (iv) development of novel software to facilitate curation and use of the database, (v) addition of new subclasses of transport systems including 11 novel types of channels and 3 types of group translocators and (vi) the inclusion of many man-made (artificial) transmembrane pores/channels and carriers.</p>]]></description>
            <pubDate><![CDATA[2020-11-10T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The ModelSEED Biochemistry Database for the integration of metabolic annotations and the reconstruction, comparison and analysis of metabolic models for plants, fungi and microbes]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609152536-f1c05abd-f917-4ab0-b571-33974aef19e9/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa746</link>
            <description><![CDATA[<p class="para" id="N65541">For over 10 years, ModelSEED has been a primary resource for the construction of draft genome-scale metabolic models based on annotated microbial or plant genomes. Now being released, the biochemistry database serves as the foundation of biochemical data underlying ModelSEED and KBase. The biochemistry database embodies several properties that, taken together, distinguish it from other published biochemistry resources by: (i) including compartmentalization, transport reactions, charged molecules and proton balancing on reactions; (ii) being extensible by the user community, with all data stored in GitHub; and (iii) design as a biochemical ‘Rosetta Stone’ to facilitate comparison and integration of annotations from many different tools and databases. The database was constructed by combining chemical data from many resources, applying standard transformations, identifying redundancies and computing thermodynamic properties. The ModelSEED biochemistry is continually tested using flux balance analysis to ensure the biochemical network is modeling-ready and capable of simulating diverse phenotypes. Ontologies can be designed to aid in comparing and reconciling metabolic reconstructions that differ in how they represent various metabolic pathways. ModelSEED now includes 33,978 compounds and 36,645 reactions, available as a set of extensible files on GitHub, and available to search at https://modelseed.org/biochem and KBase.</p>]]></description>
            <pubDate><![CDATA[2020-09-28T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[HMPDACC: a Human Microbiome Project Multi-omic data resource]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609133780-4090a526-70e7-4f6f-b669-ff6b2b1af361/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa996</link>
            <description><![CDATA[<p class="para" id="N65541">The Human Microbiome Project (HMP) explored microbial communities of the human body in both healthy and disease states. Two phases of the HMP (HMP and iHMP) together generated &gt;48TB of data (public and controlled access) from multiple, varied omics studies of both the microbiome and associated hosts. The Human Microbiome Project Data Coordination Center (HMPDACC) was established to provide a portal to access data and resources produced by the HMP. The HMPDACC provides a unified data repository, multi-faceted search functionality, analysis pipelines and standardized protocols to facilitate community use of HMP data. Recent efforts have been put toward making HMP data more findable, accessible, interoperable and reusable. HMPDACC resources are freely available at www.hmpdacc.org.</p>]]></description>
            <pubDate><![CDATA[2020-12-10T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[DescribePROT: database of amino acid-level protein structure and function predictions]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609118275-e18b7c3f-cdb0-4ffb-ab63-6aefe9953e01/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa931</link>
            <description><![CDATA[<p class="para" id="N65541">We present DescribePROT, the database of predicted amino acid-level descriptors of structure and function of proteins. DescribePROT delivers a comprehensive collection of 13 complementary descriptors predicted using 10 popular and accurate algorithms for 83 complete proteomes that cover key model organisms. The current version includes 7.8 billion predictions for close to 600 million amino acids in 1.4 million proteins. The descriptors encompass sequence conservation, position specific scoring matrix, secondary structure, solvent accessibility, intrinsic disorder, disordered linkers, signal peptides, MoRFs and interactions with proteins, DNA and RNAs. Users can search DescribePROT by the amino acid sequence and the UniProt accession number and entry name. The pre-computed results are made available instantaneously. The predictions can be accesses via an interactive graphical interface that allows simultaneous analysis of multiple descriptors and can be also downloaded in structured formats at the protein, proteome and whole database scale. The putative annotations included by DescriPROT are useful for a broad range of studies, including: investigations of protein function, applied projects focusing on therapeutics and diseases, and in the development of predictors for other protein sequence descriptors. Future releases will expand the coverage of DescribePROT. DescribePROT can be accessed at http://biomine.cs.vcu.edu/servers/DESCRIBEPROT/.</p>]]></description>
            <pubDate><![CDATA[2020-10-29T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[TransCirc: an interactive database for translatable circular RNAs based on multi-omics evidence]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609035823-e4c989a2-f172-4cfc-a14c-4bbb2596f903/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa823</link>
            <description><![CDATA[<p class="para" id="N65541">TransCirc (https://www.biosino.org/transcirc/) is a specialized database that provide comprehensive evidences supporting the translation potential of circular RNAs (circRNAs). This database was generated by integrating various direct and indirect evidences to predict coding potential of each human circRNA and the putative translation products. Seven types of evidences for circRNA translation were included: (i) ribosome/polysome binding evidences supporting the occupancy of ribosomes onto circRNAs; (ii) experimentally mapped translation initiation sites on circRNAs; (iii) internal ribosome entry site on circRNAs; (iv) published N-6-methyladenosine modification data in circRNA that promote translation initiation; (v) lengths of the circRNA specific open reading frames; (vi) sequence composition scores from a machine learning prediction of all potential open reading frames; (vii) mass spectrometry data that directly support the circRNA encoded peptides across back-splice junctions. TransCirc provides a user-friendly searching/browsing interface and independent lines of evidences to predicte how likely a circRNA can be translated. In addition, several flexible tools have been developed to aid retrieval and analysis of the data. TransCirc can serve as an important resource for investigating the translation capacity of circRNAs and the potential circRNA-encoded peptides, and can be expanded to include new evidences or additional species in the future.</p>]]></description>
            <pubDate><![CDATA[2020-10-19T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[CSVS, a crowdsourcing database of the Spanish population genetic variability]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765609000552-0aacf167-d450-4174-bb7c-bb29e5c34ca4/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa794</link>
            <description><![CDATA[<p class="para" id="N65541">The knowledge of the genetic variability of the local population is of utmost importance in personalized medicine and has been revealed as a critical factor for the discovery of new disease variants. Here, we present the Collaborative Spanish Variability Server (CSVS), which currently contains more than 2000 genomes and exomes of unrelated Spanish individuals. This database has been generated in a collaborative crowdsourcing effort collecting sequencing data produced by local genomic projects and for other purposes. Sequences have been grouped by ICD10 upper categories. A web interface allows querying the database removing one or more ICD10 categories. In this way, aggregated counts of allele frequencies of the pseudo-control Spanish population can be obtained for diseases belonging to the category removed. Interestingly, in addition to pseudo-control studies, some population studies can be made, as, for example, prevalence of pharmacogenomic variants, etc. In addition, this genomic data has been used to define the first Spanish Genome Reference Panel (SGRP1.0) for imputation. This is the first local repository of variability entirely produced by a crowdsourcing effort and constitutes an example for future initiatives to characterize local variability worldwide. CSVS is also part of the GA4GH Beacon network.</p><p class="para" id="N65543">CSVS can be accessed at: http://csvs.babelomics.org/.</p>]]></description>
            <pubDate><![CDATA[2020-09-29T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[FlyRNAi.org—the database of the Drosophila RNAi screening center and transgenic RNAi project: 2021 update]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608996346-bf036295-bedd-4ec1-9f21-aa003f6acd36/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa936</link>
            <description><![CDATA[<p class="para" id="N65541">The FlyRNAi database at the Drosophila RNAi Screening Center and Transgenic RNAi Project (DRSC/TRiP) provides a suite of online resources that facilitate functional genomics studies with a special emphasis on <i>Drosophila melanogaster</i>. Currently, the database provides: gene-centric resources that facilitate ortholog mapping and mining of information about orthologs in common genetic model species; reagent-centric resources that help researchers identify RNAi and CRISPR sgRNA reagents or designs; and data-centric resources that facilitate visualization and mining of transcriptomics data, protein modification data, protein interactions, and more. Here, we discuss updated and new features that help biological and biomedical researchers efficiently identify, visualize, analyze, and integrate information and data for <i>Drosophila</i> and other species. Together, these resources facilitate multiple steps in functional genomics workflows, from building gene and reagent lists to management, analysis, and integration of data.</p>]]></description>
            <pubDate><![CDATA[2020-10-26T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Open Targets Platform: supporting systematic drug–target identification and prioritisation]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608990890-c424c169-c7aa-4795-b8d3-270e7f632725/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1027</link>
            <description><![CDATA[<p class="para" id="N65541">The Open Targets Platform (https://www.targetvalidation.org/) provides users with a queryable knowledgebase and user interface to aid systematic target identification and prioritisation for drug discovery based upon underlying evidence. It is publicly available and the underlying code is open source. Since our last update two years ago, we have had 10 releases to maintain and continuously improve evidence for target–disease relationships from 20 different data sources. In addition, we have integrated new evidence from key datasets, including prioritised targets identified from genome-wide CRISPR knockout screens in 300 cancer models (Project Score), and GWAS/UK BioBank statistical genetic analysis evidence from the Open Targets Genetics Portal. We have evolved our evidence scoring framework to improve target identification. To aid the prioritisation of targets and inform on the potential impact of modulating a given target, we have added evaluation of post-marketing adverse drug reactions and new curated information on target tractability and safety. We have also developed the user interface and backend technologies to improve performance and usability. In this article, we describe the latest enhancements to the Platform, to address the fundamental challenge that developing effective and safe drugs is difficult and expensive.</p>]]></description>
            <pubDate><![CDATA[2020-11-16T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[TREND-DB—a transcriptome-wide atlas of the dynamic landscape of alternative polyadenylation]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608955862-55dfdeaa-267a-471d-b5e2-b955bdac581b/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa722</link>
            <description><![CDATA[<p class="para" id="N65541">Alternative polyadenylation (APA) profoundly expands the transcriptome complexity. Perturbations of APA can disrupt biological processes, ultimately resulting in devastating disorders. A major challenge in identifying mechanisms and consequences of APA (and its perturbations) lies in the complexity of RNA 3′ end processing, involving poorly conserved RNA motifs and multi-component complexes consisting of far more than 50 proteins. This is further complicated in that RNA 3′ end maturation is closely linked to transcription, RNA processing and even epigenetic (histone/DNA/RNA) modifications. Here, we present TREND-DB (http://shiny.imbei.uni-mainz.de:3838/trend-db), a resource cataloging the dynamic landscape of APA after depletion of &gt;170 proteins involved in various facets of transcriptional, co- and post-transcriptional gene regulation, epigenetic modifications and further processes. TREND-DB visualizes the dynamics of transcriptome 3′ end diversification (TREND) in a highly interactive manner; it provides a global APA network map and allows interrogating genes affected by specific APA-regulators and vice versa. It also permits condition-specific functional enrichment analyses of APA-affected genes, which suggest wide biological and clinical relevance across all RNAi conditions. The implementation of the UCSC Genome Browser provides additional customizable layers of gene regulation accounting for individual transcript isoforms (e.g. epigenetics, miRNA-binding sites and RNA-binding proteins). TREND-DB thereby fosters disentangling the role of APA for various biological programs, including potential disease mechanisms, and helps identify their diagnostic and therapeutic potential.</p>]]></description>
            <pubDate><![CDATA[2020-09-25T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[mMGE: a database for human metagenomic extrachromosomal mobile genetic elements]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608949944-66208daa-b95c-4f4a-b136-ba0715869c08/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa869</link>
            <description><![CDATA[<p class="para" id="N65541">Extrachromosomal mobile genetic elements (eMGEs), including phages and plasmids, that can move across different microbes, play important roles in genome evolution and shaping the structure of microbial communities. However, we still know very little about eMGEs, especially their abundances, distributions and putative functions in microbiomes. Thus, a comprehensive description of eMGEs is of great utility. Here we present mMGE, a comprehensive catalog of 517 251 non-redundant eMGEs, including 92 492 plasmids and 424 759 phages, derived from diverse body sites of 66 425 human metagenomic samples. About half the eMGEs could be further grouped into 70 074 clusters using relaxed criteria (referred as to eMGE clusters below). We provide extensive annotations of the identified eMGEs including sequence characteristics, taxonomy affiliation, gene contents and their prokaryotic hosts. We also calculate the prevalence, both within and across samples for each eMGE and eMGE cluster, enabling users to see putative associations of eMGEs with human phenotypes or their distribution preferences. All eMGE records can be browsed or queried in multiple ways, such as eMGE clusters, metagenomic samples and associated hosts. The mMGE is equipped with a user-friendly interface and a BLAST server, facilitating easy access/queries to all its contents easily. mMGE is freely available for academic use at: https://mgedb.comp-sysbio.org.</p>]]></description>
            <pubDate><![CDATA[2020-10-19T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[RBP2GO: a comprehensive pan-species database on RNA-binding proteins, their interactions and functions]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608944311-6f19245e-89ee-495f-aee6-457560bc8345/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1040</link>
            <description><![CDATA[<p class="para" id="N65541">RNA–protein complexes have emerged as central players in numerous key cellular processes with significant relevance in health and disease. To further deepen our knowledge of RNA-binding proteins (RBPs), multiple proteome-wide strategies have been developed to identify RBPs in different species leading to a large number of studies contributing experimentally identified as well as predicted RBP candidate catalogs. However, the rapid evolution of the field led to an accumulation of isolated datasets, hampering the access and comparison of their valuable content. Moreover, tools to link RBPs to cellular pathways and functions were lacking. Here, to facilitate the efficient screening of the RBP resources, we provide RBP2GO (https://RBP2GO.DKFZ.de), a comprehensive database of all currently available proteome-wide datasets for RBPs across 13 species from 53 studies including 105 datasets identifying altogether 22 552 RBP candidates. These are combined with the information on RBP interaction partners and on the related biological processes, molecular functions and cellular compartments. RBP2GO offers a user-friendly web interface with an RBP scoring system and powerful advanced search tools allowing forward and reverse searches connecting functions and RBPs to stimulate new research directions.</p><p class="para" id="N65542">
<div class="section" id="ga1"><div class="img"><div class="imgeVideo"><div class="img-fullscreenIcon" onClick="javascript:showImageContent('ga1');"><img src="/public/images/journalImg/fullscreen.png"/></div><div class="imageVideo"><img src="/dataresources/secured/content-1765608944311-6f19245e-89ee-495f-aee6-457560bc8345/assets/gkaa1040gra1.jpg" alt="RBP2GO (https://RBP2GO.DKFZ.de) is a comprehensive database of all currently available proteome-wide datasets for RNA-binding proteins, their functions and interaction partners.Question: is it necessary to have a caption for the graphical abstract? I cannot see them on the website or PDFs of the articles."/></div></div><div class="imgeVideoCaption" id="N65544"><div class="captionTitle">Graphical Abstract</div><div class="captionText">                                      RBP2GO (https://RBP2GO.DKFZ.de) is a comprehensive database of all currently available proteome-wide datasets for RNA-binding proteins, their functions and interaction partners.Question: is it necessary to have a caption for the graphical abstract? I cannot see them on the website or PDFs of the articles.</div></div></div></div>
</p>]]></description>
            <pubDate><![CDATA[2020-11-16T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[COVID19 Drug Repository: text-mining the literature in search of putative COVID19 therapeutics]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608938800-ed2e9ac9-e388-4121-9467-468055a684bb/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa969</link>
            <description><![CDATA[<p class="para" id="N65541">The recent outbreak of COVID-19 has generated an enormous amount of Big Data. To date, the COVID-19 Open Research Dataset (CORD-19), lists ∼130,000 articles from the WHO COVID-19 database, PubMed Central, medRxiv, and bioRxiv, as collected by Semantic Scholar. According to LitCovid (11 August 2020), ∼40,300 COVID19-related articles are currently listed in PubMed. It has been shown in clinical settings that the analysis of past research results and the mining of available data can provide novel opportunities for the successful application of currently approved therapeutics and their combinations for the treatment of conditions caused by a novel SARS-CoV-2 infection. As such, effective responses to the pandemic require the development of efficient applications, methods and algorithms for data navigation, text-mining, clustering, classification, analysis, and reasoning. Thus, our COVID19 Drug Repository represents a modular platform for drug data navigation and analysis, with an emphasis on COVID-19-related information currently being reported. The COVID19 Drug Repository enables users to focus on different levels of complexity, starting from general information about (FDA-) approved drugs, PubMed references, clinical trials, recipes as well as the descriptions of molecular mechanisms of drugs’ action. Our COVID19 drug repository provide a most updated world-wide collection of drugs that has been repurposed for COVID19 treatments around the world.</p>]]></description>
            <pubDate><![CDATA[2020-11-09T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[DockCoV2: a drug database against SARS-CoV-2]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608922070-595be358-2b22-4b2b-8cd2-dcfc788e4cd1/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa861</link>
            <description><![CDATA[<p class="para" id="N65541">The current state of the COVID-19 pandemic is a global health crisis. To fight the novel coronavirus, one of the best-known ways is to block enzymes essential for virus replication. Currently, we know that the SARS-CoV-2 virus encodes about 29 proteins such as spike protein, 3C-like protease (3CLpro), RNA-dependent RNA polymerase (RdRp), Papain-like protease (PLpro), and nucleocapsid (N) protein. SARS-CoV-2 uses human angiotensin-converting enzyme 2 (ACE2) for viral entry and transmembrane serine protease family member II (TMPRSS2) for spike protein priming. Thus in order to speed up the discovery of potential drugs, we develop DockCoV2, a drug database for SARS-CoV-2. DockCoV2 focuses on predicting the binding affinity of FDA-approved and Taiwan National Health Insurance (NHI) drugs with the seven proteins mentioned above. This database contains a total of 3,109 drugs. DockCoV2 is easy to use and search against, is well cross-linked to external databases, and provides the state-of-the-art prediction results in one site. Users can download their drug-protein docking data of interest and examine additional drug-related information on DockCoV2. Furthermore, DockCoV2 provides experimental information to help users understand which drugs have already been reported to be effective against MERS or SARS-CoV. DockCoV2 is available at https://covirus.cc/drugs/.</p>]]></description>
            <pubDate><![CDATA[2020-10-09T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[PROTAC-DB: an online database of PROTACs]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608917647-008e9025-ffba-4ae5-8952-f8b6f46a4f7d/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa807</link>
            <description><![CDATA[<p class="para" id="N65541">Proteolysis-targeting chimeras (PROTACs), which selectively degrade targeted proteins by the ubiquitin-proteasome system, have emerged as a novel therapeutic technology with potential advantages over traditional inhibition strategies. In the past few years, this technology has achieved substantial progress and two PROTACs have been advanced into phase I clinical trials. However, this technology is still maturing and the design of PROTACs remains a great challenge. In order to promote the rational design of PROTACs, we present PROTAC-DB, a web-based open-access database that integrates structural information and experimental data of PROTACs. Currently, PROTAC-DB consists of 1662 PROTACs, 202 warheads (small molecules that target the proteins of interest), 65 E3 ligands (small molecules capable of recruiting E3 ligases) and 806 linkers, as well as their chemical structures, biological activities, and physicochemical properties. Except the biological activities of warheads and E3 ligands, PROTAC-DB also provides the degradation capacities, binding affinities and cellular activities for PROTACs. PROTAC-DB can be queried with two general searching approaches: text-based (target name, compound name or ID) and structure-based. In addition, for the convenience of users, a filtering tool for the searching results based on the physicochemical properties of compounds is also offered. PROTAC-DB is freely accessible at http://cadd.zju.edu.cn/protacdb/.</p><p class="para" id="N65542">
<div class="section" id="ga1"><div class="img"><div class="imgeVideo"><div class="img-fullscreenIcon" onClick="javascript:showImageContent('ga1');"><img src="/public/images/journalImg/fullscreen.png"/></div><div class="imageVideo"><img src="/dataresources/secured/content-1765608917647-008e9025-ffba-4ae5-8952-f8b6f46a4f7d/assets/gkaa807gra1.jpg" alt="The three basic steps to construct PROTAC-DB: (I) literature searching, (II) data collection, and (III) data processing by separating PROTACs into warheads, E3 ligands and linkers."/></div></div><div class="imgeVideoCaption" id="N65544"><div class="captionTitle">Graphical Abstract</div><div class="captionText">                                      The three basic steps to construct PROTAC-DB: (I) literature searching, (II) data collection, and (III) data processing by separating PROTACs into warheads, E3 ligands and linkers.</div></div></div></div>
</p>]]></description>
            <pubDate><![CDATA[2020-10-03T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Genenames.org: the HGNC and VGNC resources in 2021]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608904045-2da0512f-cada-4df9-bb35-cc762b8a2a9d/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa980</link>
            <description><![CDATA[<p class="para" id="N65541">The HUGO Gene Nomenclature Committee (HGNC) based at EMBL’s European Bioinformatics Institute (EMBL-EBI) assigns unique symbols and names to human genes. There are over 42,000 approved gene symbols in our current database of which over 19 000 are for protein-coding genes. While we still update placeholder and problematic symbols, we are working towards stabilizing symbols where possible; over 2000 symbols for disease associated genes are now marked as stable in our symbol reports. All of our data is available at the HGNC website https://www.genenames.org. The Vertebrate Gene Nomenclature Committee (VGNC) was established to assign standardized nomenclature in line with human for vertebrate species lacking their own nomenclature committee. In addition to the previous VGNC core species of chimpanzee, cow, horse and dog, we now name genes in cat, macaque and pig. Gene groups have been added to VGNC and currently include two complex families: olfactory receptors (ORs) and cytochrome P450s (CYPs). In collaboration with specialists we have also named CYPs in species beyond our core set. All VGNC data is available at https://vertebrate.genenames.org/. This article provides an overview of our online data and resources, focusing on updates over the last two years.</p>]]></description>
            <pubDate><![CDATA[2020-11-05T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[PINA 3.0: mining cancer interactome]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608899841-76bbecee-d98b-4663-bae8-769185f35ed3/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1075</link>
            <description><![CDATA[<p class="para" id="N65541">Protein–protein interactions (PPIs) are crucial to mediate biological functions, and understanding PPIs in cancer type-specific context could help decipher the underlying molecular mechanisms of tumorigenesis and identify potential therapeutic options. Therefore, we update the Protein Interaction Network Analysis (PINA) platform to version 3.0, to integrate the unified human interactome with RNA-seq transcriptomes and mass spectrometry-based proteomes across tens of cancer types. A number of new analytical utilities were developed to help characterize the cancer context for a PPI network, which includes inferring proteins with expression specificity and identifying candidate prognosis biomarkers, putative cancer drivers, and therapeutic targets for a specific cancer type; as well as identifying pairs of co-expressing interacting proteins across cancer types. Furthermore, a brand-new web interface has been designed to integrate these new utilities within an interactive network visualization environment, which allows users to quickly and comprehensively investigate the roles of human interacting proteins in a cancer type-specific context. PINA is freely available at https://omics.bjcancer.org/pina/.</p><p class="para" id="N65542">
<div class="section" id="ga1"><div class="img"><div class="imgeVideo"><div class="img-fullscreenIcon" onClick="javascript:showImageContent('ga1');"><img src="/public/images/journalImg/fullscreen.png"/></div><div class="imageVideo"><img src="/dataresources/secured/content-1765608899841-76bbecee-d98b-4663-bae8-769185f35ed3/assets/gkaa1075gra1.jpg" alt="A schematic overview of PINA 3.0."/></div></div><div class="imgeVideoCaption" id="N65544"><div class="captionTitle">Graphical Abstract</div><div class="captionText">                                      A schematic overview of PINA 3.0.</div></div></div></div>
</p>]]></description>
            <pubDate><![CDATA[2020-11-24T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[ConjuPepDB: a database of peptide–drug conjugates]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608845247-9aa45b30-45c7-4a76-9d5a-c95dac44541f/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa950</link>
            <description><![CDATA[<p class="para" id="N65541">Peptide–drug conjugates are organic molecules composed of (i) a small drug molecule, (ii) a peptide and (iii) a linker. The drug molecule is mandatory for the biological action, however, its efficacy can be enhanced by targeted delivery, which often also reduces unwanted side effects. For site-specificity the peptide part is mainly responsible. The linker attaches chemically the drug to the peptide, but it could also be biodegradable which ensures controlled liberation of the small drug. Despite the importance of the field, there is no public comprehensive database on these species. Herein we describe ConjuPepBD, a freely available, fully annotated and manually curated database of peptide drug conjugates. ConjuPepDB contains basic information about the entries, e.g. CAS number. Furthermore, it also implies their biomedical application and the type of chemical conjugation employed. It covers more than 1600 conjugates from ∼230 publications. The web-interface is user-friendly, intuitive, and useable on several devices, e.g. phones, tablets, PCs. The webpage allows the user to search for content using numerous criteria, chemical structure and a help page is also provided. Besides giving quick insight for newcomers, ConjuPepDB is hoped to be also helpful for researchers from various related fields. The database is accessible at: https://conjupepdb.ttk.hu/.</p><p class="para" id="N65542">
<div class="section" id="ga1"><div class="img"><div class="imgeVideo"><div class="img-fullscreenIcon" onClick="javascript:showImageContent('ga1');"><img src="/public/images/journalImg/fullscreen.png"/></div><div class="imageVideo"><img src="/dataresources/secured/content-1765608845247-9aa45b30-45c7-4a76-9d5a-c95dac44541f/assets/gkaa950gra1.jpg" alt="Peptide drug conjugates as novel class of potential drug candidates are collected."/></div></div><div class="imgeVideoCaption" id="N65544"><div class="captionTitle">Graphical Abstract</div><div class="captionText">                                      Peptide drug conjugates as novel class of potential drug candidates are collected.</div></div></div></div>
</p>]]></description>
            <pubDate><![CDATA[2020-10-30T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Genome Variation Map: a worldwide collection of genome variations across multiple species]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608841275-1971c587-5fb3-4795-af9e-fb3ad106c891/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1005</link>
            <description><![CDATA[<p class="para" id="N65541">The Genome Variation Map (GVM; http://bigd.big.ac.cn/gvm/) is a public data repository of genome variations. It aims to collect and integrate genome variations for a wide range of species, accepts submissions of different variation types from all over the world and provides free open access to all publicly available data in support of worldwide research activities. Compared with the previous version, particularly, a total of 22 species, 115 projects, 55 935 samples, 463 429 609 variants, 66 220 associations and 56 submissions (as of 7 September 2020) were newly added in the current version of GVM. In the current release, GVM houses a total of ∼960 million variants from 41 species, including 13 animals, 25 plants and 3 viruses. Moreover, it incorporates 64 819 individual genotypes and 260 393 manually curated high-quality genotype-to-phenotype associations. Since its inception, GVM has archived genomic variation data of 43 754 samples submitted by worldwide users and served &gt;1 million data download requests. Collectively, as a core resource in the National Genomics Data Center, GVM provides valuable genome variations for a diversity of species and thus plays an important role in both functional genomics studies and molecular breeding.</p>]]></description>
            <pubDate><![CDATA[2020-11-10T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[BiG-FAM: the biosynthetic gene cluster families database]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608792703-23e69fa1-30a1-4f5f-8f43-fb7650141ade/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa812</link>
            <description><![CDATA[<p class="para" id="N65541">Computational analysis of biosynthetic gene clusters (BGCs) has revolutionized natural product discovery by enabling the rapid investigation of secondary metabolic potential within microbial genome sequences. Grouping homologous BGCs into Gene Cluster Families (GCFs) facilitates mapping their architectural and taxonomic diversity and provides insights into the novelty of putative BGCs, through dereplication with BGCs of known function. While multiple databases exist for exploring BGCs from publicly available data, no public resources exist that focus on GCF relationships. Here, we present BiG-FAM, a database of 29,955 GCFs capturing the global diversity of 1,225,071 BGCs predicted from 209,206 publicly available microbial genomes and metagenome-assembled genomes (MAGs). The database offers rich functionalities, such as multi-criterion GCF searches, direct links to BGC databases such as antiSMASH-DB, and rapid GCF annotation of user-supplied BGCs from antiSMASH results. BiG-FAM can be accessed online at https://bigfam.bioinformatics.nl.</p><p class="para" id="N65542">
<div class="section" id="ga1"><div class="img"><div class="imgeVideo"><div class="img-fullscreenIcon" onClick="javascript:showImageContent('ga1');"><img src="/public/images/journalImg/fullscreen.png"/></div><div class="imageVideo"><img src="/dataresources/secured/content-1765608792703-23e69fa1-30a1-4f5f-8f43-fb7650141ade/assets/gkaa812gra1.jpg" alt="Constructed based on a large-scale homology analysis of 1.2 million biosynthetic gene clusters, BiG-FAM provides a platform to explore their genomic diversity and to discover their relationships to newly sequenced ones."/></div></div><div class="imgeVideoCaption" id="N65544"><div class="captionTitle">Graphical Abstract</div><div class="captionText">                                      Constructed based on a large-scale homology analysis of 1.2 million biosynthetic gene clusters, BiG-FAM provides a platform to explore their genomic diversity and to discover their relationships to newly sequenced ones.</div></div></div></div>
</p>]]></description>
            <pubDate><![CDATA[2020-10-03T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[EVLncRNAs 2.0: an updated database of manually curated functional long non-coding RNAs validated by low-throughput experiments]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608764401-f610cb2f-757f-431d-a6bc-e93441237197/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1076</link>
            <description><![CDATA[<p class="para" id="N65541">Long non-coding RNAs (lncRNAs) play important functional roles in many diverse biological processes. However, not all expressed lncRNAs are functional. Thus, it is necessary to manually collect all experimentally validated functional lncRNAs (EVlncRNA) with their sequences, structures, and functions annotated in a central database. The first release of such a database (EVLncRNAs) was made using the literature prior to 1 May 2016. Since then (till 15 May 2020), 19 245 articles related to lncRNAs have been published. In EVLncRNAs 2.0, these articles were manually examined for a major expansion of the data collected. Specifically, the number of annotated EVlncRNAs, associated diseases, lncRNA-disease associations, and interaction records were increased by 260%, 320%, 484% and 537%, respectively. Moreover, the database has added several new categories: 8 lncRNA structures, 33 exosomal lncRNAs, 188 circular RNAs, and 1079 drug-resistant, chemoresistant, and stress-resistant lncRNAs. All records have checked against known retraction and fake articles. This release also comes with a highly interactive visual interaction network that facilitates users to track the underlying relations among lncRNAs, miRNAs, proteins, genes and other functional elements. Furthermore, it provides links to four new bioinformatics tools with improved data browsing and searching functionality. EVLncRNAs 2.0 is freely available at https://www.sdklab-biophysics-dzu.net/EVLncRNAs2/.</p>]]></description>
            <pubDate><![CDATA[2020-11-22T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[OMA orthology in 2021: website overhaul, conserved isoforms, ancestral gene order and more]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608718197-4ab4ff15-3fb3-48e4-992f-52512303079f/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1007</link>
            <description><![CDATA[<p class="para" id="N65541">OMA is an established resource to elucidate evolutionary relationships among genes from currently 2326 genomes covering all domains of life. OMA provides pairwise and groupwise orthologs, functional annotations, local and global gene order conservation (synteny) information, among many other functions. This update paper describes the reorganisation of the database into gene-, group- and genome-centric pages. Other new and improved features are detailed, such as reporting of the evolutionarily best conserved isoforms of alternatively spliced genes, the inferred local order of ancestral genes, phylogenetic profiling, better cross-references, fast genome mapping, semantic data sharing via RDF, as well as a special coronavirus OMA with 119 viruses from the Nidovirales order, including SARS-CoV-2, the agent of the COVID-19 pandemic. We conclude with improvements to the documentation of the resource through primers, tutorials and short videos. OMA is accessible at https://omabrowser.org.</p>]]></description>
            <pubDate><![CDATA[2020-11-11T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[TBDB: a database of structurally annotated T-box riboswitch:tRNA pairs]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608674614-a1a917fc-54b2-4c45-bb23-806cc54ff715/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa721</link>
            <description><![CDATA[<p class="para" id="N65541">T-box riboswitches constitute a large family of tRNA-binding leader sequences that play a central role in gene regulation in many gram-positive bacteria. Accurate inference of the tRNA binding to T-box riboswitches is critical to predict their cis-regulatory activity. However, there is no central repository of information on the tRNA binding specificities of T-box riboswitches, and <i>de novo</i> prediction of binding specificities requires advanced knowledge of computational tools to annotate riboswitch secondary structure features. Here, we present the T-box Riboswitch Annotation Database (TBDB, https://tbdb.io), an open-access database with a collection of 23,535 T-box riboswitch sequences, spanning the major phyla of 3,632 bacterial species. Among structural predictions, the TBDB also identifies specifier sequences, cognate tRNA binding partners, and downstream regulatory targets. To our knowledge, the TBDB presents the largest collection of feature, sequence, and structural annotations carried out on this important family of regulatory RNA.</p>]]></description>
            <pubDate><![CDATA[2020-09-03T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[PLncDB V2.0: a comprehensive encyclopedia of plant long noncoding RNAs]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608669943-1cc5351a-4253-4d63-a079-8a13e776d005/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa910</link>
            <description><![CDATA[<p class="para" id="N65541">Long noncoding RNAs (lncRNAs) are transcripts longer than 200 nucleotides with little or no protein coding potential. The expanding list of lncRNAs and accumulating evidence of their functions in plants have necessitated the creation of a comprehensive database for lncRNA research. However, currently available plant lncRNA databases have some deficiencies, including the lack of lncRNA data from some model plants, uneven annotation standards, a lack of visualization for expression patterns, and the absence of epigenetic information. To overcome these problems, we upgraded our Plant Long noncoding RNA Database (PLncDB, http://plncdb.tobaccodb.org/), which was based on a uniform annotation pipeline. PLncDB V2.0 currently contains 1 246 372 lncRNAs for 80 plant species based on 13 834 RNA-Seq datasets, integrating lncRNA information from four other resources including EVLncRNAs, RNAcentral and etc. Expression patterns and epigenetic signals can be visualized using multiple tools (JBrowse, eFP Browser and EPexplorer). Targets and regulatory networks for lncRNAs are also provided for function exploration. In addition, PLncDB V2.0 is hierarchical and user-friendly and has five built-in search engines. We believe PLncDB V2.0 is useful for the plant lncRNA community and data mining studies and provides a comprehensive resource for data-driven lncRNA research in plants.</p>]]></description>
            <pubDate><![CDATA[2020-10-20T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[LincSNP 3.0: an updated database for linking functional variants to human long non-coding RNAs, circular RNAs and their regulatory elements]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608612253-2c66ec92-7d91-4e15-9e6a-7aad7b883592/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1037</link>
            <description><![CDATA[<p class="para" id="N65541">We describe an updated comprehensive database, LincSNP 3.0 (http://bioinfo.hrbmu.edu.cn/LincSNP), which aims to document and annotate disease or phenotype-associated variants in human long non-coding RNAs (lncRNAs) and circular RNAs (circRNAs) or their regulatory elements. LincSNP 3.0 has updated with several novel features, including (i) more types of variants including single nucleotide polymorphisms (SNPs), linkage disequilibrium SNPs (LD SNPs), somatic mutations and RNA editing sites have been expanded; (ii) more regulatory elements including transcription factor binding sites (TFBSs), enhancers, DNase I hypersensitive sites (DHSs), topologically associated domains (TADs), footprintss, methylations and open chromatin regions have been added; (iii) the associations among circRNAs, regulatory elements and variants have been identified; (iv) more experimentally supported variant-lncRNA/circRNA-disease/phenotype associations have been manually collected; (v) the sources of lncRNAs, circRNAs, SNPs, somatic mutations and RNA editing sites have been updated. Moreover, four flexible online tools including Genome Browser, Variant Mapper, Circos Plotter and Functional Annotation have been developed to retrieve, visualize and analyze the data. Collectively, LincSNP 3.0 provides associations among functional variants, regulatory elements, lncRNAs and circRNAs in diseases. It will serve as an important and continually updated resource for investigating functions and mechanisms of lncRNAs and circRNAs in diseases.</p>]]></description>
            <pubDate><![CDATA[2020-11-21T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[dbCAN-PUL: a database of experimentally characterized CAZyme gene clusters and their substrates]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608593483-2faafdd7-8f7f-4df3-aa7c-8f296cf0ba62/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa742</link>
            <description><![CDATA[<p class="para" id="N65541">PULs (<span style="text-decoration: underline">p</span>olysaccharide <span style="text-decoration: underline">u</span>tilization <span style="text-decoration: underline">l</span>oci) are discrete gene clusters of CAZymes (<span style="text-decoration: underline">C</span>arbohydrate <span style="text-decoration: underline">A</span>ctive En<span style="text-decoration: underline">Zymes</span>) and other genes that work together to digest and utilize carbohydrate substrates. While PULs have been extensively characterized in <i>Bacteroidetes</i>, there exist PULs from other bacterial phyla, as well as archaea and metagenomes, that remain to be catalogued in a database for efficient retrieval. We have developed an online database dbCAN-PUL (http://bcb.unl.edu/dbCAN_PUL/) to display experimentally verified CAZyme-containing PULs from literature with pertinent metadata, sequences, and annotation. Compared to other online CAZyme and PUL resources, dbCAN-PUL has the following new features: (i) Batch download of PUL data by target substrate, species/genome, genus, or experimental characterization method; (ii) Annotation for each PUL that displays associated metadata such as substrate(s), experimental characterization method(s) and protein sequence information, (iii) Links to external annotation pages for CAZymes (CAZy), transporters (UniProt) and other genes, (iv) Display of homologous gene clusters in GenBank sequences via integrated MultiGeneBlast tool and (v) An integrated BLASTX service available for users to query their sequences against PUL proteins in dbCAN-PUL. With these features, dbCAN-PUL will be an important repository for CAZyme and PUL research, complementing our other web servers and databases (dbCAN2, dbCAN-seq).</p>]]></description>
            <pubDate><![CDATA[2020-09-17T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[MetaNetX/MNXref: unified namespace for metabolites and biochemical reactions in the context of metabolic models]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608586010-9863c858-5162-495b-9d41-3c76b2057aea/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa992</link>
            <description><![CDATA[<p class="para" id="N65541">MetaNetX/MNXref is a reconciliation of metabolites and biochemical reactions providing cross-links between major public biochemistry and Genome-Scale Metabolic Network (GSMN) databases. The new release brings several improvements with respect to the quality of the reconciliation, with particular attention dedicated to preserving the intrinsic properties of GSMN models. The MetaNetX website (https://www.metanetx.org/) provides access to the full database and online services. A major improvement is for mapping of user-provided GSMNs to MXNref, which now provides diagnostic messages about model content. In addition to the website and flat files, the resource can now be accessed through a SPARQL endpoint (https://rdf.metanetx.org).</p>]]></description>
            <pubDate><![CDATA[2020-11-06T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[IndiGenomes: a comprehensive resource of genetic variants from over 1000 Indian genomes]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608578997-e3132a42-a5a6-41a4-8e8e-f3ab0657bb2e/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa923</link>
            <description><![CDATA[<p class="para" id="N65541">With the advent of next-generation sequencing, large-scale initiatives for mining whole genomes and exomes have been employed to better understand global or population-level genetic architecture. India encompasses more than 17% of the world population with extensive genetic diversity, but is under-represented in the global sequencing datasets. This gave us the impetus to perform and analyze the whole genome sequencing of 1029 healthy Indian individuals under the pilot phase of the ‘IndiGen’ program. We generated a compendium of 55,898,122 single allelic genetic variants from geographically distinct Indian genomes and calculated the allele frequency, allele count, allele number, along with the number of heterozygous or homozygous individuals. In the present study, these variants were systematically annotated using publicly available population databases and can be accessed through a browsable online database named as ‘IndiGenomes’ http://clingen.igib.res.in/indigen/. The IndiGenomes database will help clinicians and researchers in exploring the genetic component underlying medical conditions. Till date, this is the most comprehensive genetic variant resource for the Indian population and is made freely available for academic utility. The resource has also been accessed extensively by the worldwide community since it's launch.</p><p class="para" id="N65542">
<div class="section" id="ga1"><div class="img"><div class="imgeVideo"><div class="img-fullscreenIcon" onClick="javascript:showImageContent('ga1');"><img src="/public/images/journalImg/fullscreen.png"/></div><div class="imageVideo"><img src="/dataresources/secured/content-1765608578997-e3132a42-a5a6-41a4-8e8e-f3ab0657bb2e/assets/gkaa923gra1.jpg" alt="Overview of the Indigenomes sample collection and data analysis."/></div></div><div class="imgeVideoCaption" id="N65544"><div class="captionTitle">Graphical Abstract</div><div class="captionText">                                      Overview of the Indigenomes sample collection and data analysis.</div></div></div></div>
</p>]]></description>
            <pubDate><![CDATA[2020-10-23T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Chewie Nomenclature Server (chewie-NS): a deployable nomenclature server for easy sharing of core and whole genome MLST schemas]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608555511-62d52a6e-fb2c-4af7-9748-6587e4112c60/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa889</link>
            <description><![CDATA[<p class="para" id="N65541">Chewie Nomenclature Server (chewie-NS, https://chewbbaca.online/) allows users to share genome-based gene-by-gene typing schemas and to maintain a common nomenclature, simplifying the comparison of results. The combination between local analyses and a public repository of allelic data strikes a balance between potential confidentiality issues and the need to compare results. The possibility of deploying private instances of chewie-NS facilitates the creation of nomenclature servers with a restricted user base to allow compliance with the strictest data policies. Chewie-NS allows users to easily share their own schemas and to explore publicly available schemas, including informative statistics on schemas and loci presented in interactive charts and tables. Users can retrieve all the information necessary to run a schema locally or all the alleles identified at a particular locus. The integration with the chewBBACA suite enables users to directly upload new schemas to chewie-NS, download existing schemas and synchronize local and remote schemas from chewBBACA command line version, allowing an easier integration into high-throughput analysis pipelines. The same REST API linking chewie-NS and the chewBBACA suite supports the interaction of other interfaces or pipelines with the databases available at chewie-NS, facilitating the reusability of the stored data.</p>]]></description>
            <pubDate><![CDATA[2020-10-17T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Ensembl 2021]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608532315-5a6404d2-d449-4a36-8c8c-a35ac5d5e8de/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa942</link>
            <description><![CDATA[<p class="para" id="N65541">The Ensembl project (https://www.ensembl.org) annotates genomes and disseminates genomic data for vertebrate species. We create detailed and comprehensive annotation of gene structures, regulatory elements and variants, and enable comparative genomics by inferring the evolutionary history of genes and genomes. Our integrated genomic data are made available in a variety of ways, including genome browsers, search interfaces, specialist tools such as the Ensembl Variant Effect Predictor, download files and programmatic interfaces. Here, we present recent Ensembl developments including two new website portals. Ensembl Rapid Release (http://rapid.ensembl.org) is designed to provide core tools and services for genomes as soon as possible and has been deployed to support large biodiversity sequencing projects. Our SARS-CoV-2 genome browser (https://covid-19.ensembl.org) integrates our own annotation with publicly available genomic data from numerous sources to facilitate the use of genomics in the international scientific response to the COVID-19 pandemic. We also report on other updates to our annotation resources, tools and services. All Ensembl data and software are freely available without restriction.</p>]]></description>
            <pubDate><![CDATA[2020-11-02T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Clinically relevant updates of the HbVar database of human hemoglobin variants and thalassemia mutations]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608519687-69687228-947f-420e-a6d2-c76b351d7eb0/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa959</link>
            <description><![CDATA[<p class="para" id="N65541">HbVar (http://globin.bx.psu.edu/hbvar) is a widely-used locus-specific database (LSDB) launched 20 years ago by a multi-center academic effort to provide timely information on the numerous genomic variants leading to hemoglobin variants and all types of thalassemia and hemoglobinopathies. Here, we report several advances for the database. We made clinically relevant updates of HbVar, implemented as additional querying options in the HbVar query page, allowing the user to explore the clinical phenotype of compound heterozygous patients. We also made significant improvements to the HbVar front page, making comparative data querying, analysis and output more user-friendly. We continued to expand and enrich the regular data content, involving 1820 variants, 230 of which are new entries. We also increased the querying potential and expanded the usefulness of HbVar database in the clinical setting. These several additions, expansions and updates should improve the utility of HbVar both for the globin research community and in a clinical setting.</p>]]></description>
            <pubDate><![CDATA[2020-10-30T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[FireProt<sup>DB</sup>: database of manually curated protein stability data]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608389124-5cdd82af-878e-4aa1-ab30-231e45251b62/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa981</link>
            <description><![CDATA[<p class="para" id="N65541">The majority of naturally occurring proteins have evolved to function under mild conditions inside the living organisms. One of the critical obstacles for the use of proteins in biotechnological applications is their insufficient stability at elevated temperatures or in the presence of salts. Since experimental screening for stabilizing mutations is typically laborious and expensive, <i>in silico</i> predictors are often used for narrowing down the mutational landscape. The recent advances in machine learning and artificial intelligence further facilitate the development of such computational tools. However, the accuracy of these predictors strongly depends on the quality and amount of data used for training and testing, which have often been reported as the current bottleneck of the approach. To address this problem, we present a novel database of experimental thermostability data for single-point mutants FireProt<sup>DB</sup>. The database combines the published datasets, data extracted manually from the recent literature, and the data collected in our laboratory. Its user interface is designed to facilitate both types of the expected use: (i) the interactive explorations of individual entries on the level of a protein or mutation and (ii) the construction of highly customized and machine learning-friendly datasets using advanced searching and filtering. The database is freely available at https://loschmidt.chemi.muni.cz/fireprotdb.</p>]]></description>
            <pubDate><![CDATA[2020-11-09T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[CancerImmunityQTL: a database to systematically evaluate the impact of genetic variants on immune infiltration in human cancer]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608384474-2d076930-0a99-4ae5-a8b9-dfa7af496fad/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa805</link>
            <description><![CDATA[<p class="para" id="N65541">Tumor-infiltrating immune cells as integral component of the tumor microenvironment are associated with tumor progress, prognosis and responses to immunotherapy. Genetic variants have been demonstrated to impact tumor-infiltrating, underscoring the heritable character of immune landscape. Therefore, identification of immunity quantitative trait loci (immunQTLs), which evaluate the effect of genetic variants on immune cells infiltration, might present a critical step toward fully understanding the contribution of genetic variants in tumor development. Although emerging studies have demonstrated the determinants of germline variants on immune infiltration, no database has yet been developed to systematically analyze immunQTLs across multiple cancer types. Using genotype data from TCGA database and immune cell fractions estimated by CIBERSORT, we developed a computational pipeline to identify immunQTLs in 33 cancer types. A total of 913 immunQTLs across different cancer types were identified. Among them, 5 immunQTLs are associated with patient overall survival. Furthermore, by integrating immunQTLs with GWAS data, we identified 527 immunQTLs overlapping with known GWAS linkage disequilibrium regions. Finally, we constructed a user-friendly database, CancerImmunityQTL (http://www.cancerimmunityqtl-hust.com/) for users to browse, search and download data of interest. This database provides an informative resource to understand the germline determinants of immune infiltration in human cancer and benefit from personalized cancer immunotherapy.</p>]]></description>
            <pubDate><![CDATA[2020-10-03T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Comparative Toxicogenomics Database (CTD): update 2021]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608344973-c9ccf5bf-d468-452e-94a3-064e7de33c2b/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa891</link>
            <description><![CDATA[<p class="para" id="N65541">The public Comparative Toxicogenomics Database (CTD; http://ctdbase.org/) is an innovative digital ecosystem that relates toxicological information for chemicals, genes, phenotypes, diseases, and exposures to advance understanding about human health. Literature-based, manually curated interactions are integrated to create a knowledgebase that harmonizes cross-species heterogeneous data for chemical exposures and their biological repercussions. In this biennial update, we report a 20% increase in CTD curated content and now provide 45 million toxicogenomic relationships for over 16 300 chemicals, 51 300 genes, 5500 phenotypes, 7200 diseases and 163 000 exposure events, from 600 comparative species. Furthermore, we increase the functionality of chemical–phenotype content with new data-tabs on CTD Disease pages (to help fill in knowledge gaps for environmental health) and new phenotype search parameters (for Batch Query and Venn analysis tools). As well, we introduce new CTD Anatomy pages that allow users to uniquely explore and analyze chemical–phenotype interactions from an anatomical perspective. Finally, we have enhanced CTD Chemical pages with new literature-based chemical synonyms (to improve querying) and added 1600 amino acid-based compounds (to increase chemical landscape). Together, these updates continue to augment CTD as a powerful resource for generating testable hypotheses about the etiologies and molecular mechanisms underlying environmentally influenced diseases.</p>]]></description>
            <pubDate><![CDATA[2020-10-17T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Integration of the Drug–Gene Interaction Database (DGIdb 4.0) with open crowdsource efforts]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608332363-dd42c57a-98dd-4bae-9faf-7ce574b86f5b/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1084</link>
            <description><![CDATA[<p class="para" id="N65541">The Drug-Gene Interaction Database (DGIdb, www.dgidb.org) is a web resource that provides information on drug-gene interactions and druggable genes from publications, databases, and other web-based sources. Drug, gene, and interaction data are normalized and merged into conceptual groups. The information contained in this resource is available to users through a straightforward search interface, an application programming interface (API), and TSV data downloads. DGIdb 4.0 is the latest major version release of this database. A primary focus of this update was integration with crowdsourced efforts, leveraging the Drug Target Commons for community-contributed interaction data, Wikidata to facilitate term normalization, and export to NDEx for drug-gene interaction network representations. Seven new sources have been added since the last major version release, bringing the total number of sources included to 41. Of the previously aggregated sources, 15 have been updated. DGIdb 4.0 also includes improvements to the process of drug normalization and grouping of imported sources. Other notable updates include the introduction of a more sophisticated Query Score for interaction search results, an updated Interaction Score, the inclusion of interaction directionality, and several additional improvements to search features, data releases, licensing documentation and the application framework.</p>]]></description>
            <pubDate><![CDATA[2020-11-25T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[tsRBase: a comprehensive database for expression and function of tsRNAs in multiple species]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608319817-33f06d09-584c-4a4b-ab7d-ccd10338f68b/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa888</link>
            <description><![CDATA[<p class="para" id="N65541">tRNA-derived small RNAs (tsRNAs) are a class of novel small RNAs, ubiquitously present in prokaryotes and eukaryotes. It has been reported that tsRNAs exhibit spatiotemporal expression patterns and can function as regulatory molecules in many biological processes. Current tsRNA databases only cover limited organisms and ignore tsRNA functional characteristics. Thus, integrating more relevant tsRNA information is helpful for further exploration. Here, we present a tsRNA database, named tsRBase, which integrates the expression pattern and functional information of tsRNAs in multiple species. In tsRBase, we identified 121 942 tsRNAs by analyzing more than 14 000 publicly available small RNA-seq data covering 20 species. This database collects samples from different tissues/cell-lines, or under different treatments and genetic backgrounds, thus helps depict specific expression patterns of tsRNAs under different conditions. Importantly, to enrich our understanding of biological significance, we collected tsRNAs experimentally validated from published literatures, obtained protein-binding tsRNAs from CLIP/RIP-seq data, and identified targets of tsRNAs from CLASH and CLEAR-CLIP data. Taken together, tsRBase is the most comprehensive and systematic tsRNA repository, exhibiting all-inclusive information of tsRNAs from diverse data sources of multiple species. tsRBase is freely available at http://www.tsrbase.org.</p>]]></description>
            <pubDate><![CDATA[2020-10-17T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[cncRNAdb: a manually curated resource of experimentally supported RNAs with both protein-coding and noncoding function]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608256716-e0adf79e-adb1-40c2-ac87-7f45f2905feb/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa791</link>
            <description><![CDATA[<p class="para" id="N65541">RNA endowed with both protein-coding and noncoding functions is referred to as ‘dual-function RNA’, ‘binary functional RNA (bifunctional RNA)’ or ‘cncRNA (coding and noncoding RNA)’. Recently, an increasing number of cncRNAs have been identified, including both translated ncRNAs (ncRNAs with coding functions) and untranslated mRNAs (mRNAs with noncoding functions). However, an appropriate database for storing and organizing cncRNAs is still lacking. Here, we developed cncRNAdb, a manually curated database of experimentally supported cncRNAs, which aims to provide a resource for efficient manipulation, browsing and analysis of cncRNAs. The current version of cncRNAdb documents about 2600 manually curated entries of cncRNA functions with experimental evidence, involving more than 2,000 RNAs (including over 1300 translated ncRNAs and over 600 untranslated mRNAs) across over 20 species. In summary, we believe that cncRNAdb will help elucidate the functions and mechanisms of cncRNAs and develop new prediction methods. The database is available at http://www.rna-society.org/cncrnadb/.</p>]]></description>
            <pubDate><![CDATA[2020-10-03T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[From ArrayExpress to BioStudies]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608225115-9ae70b7a-cbf4-42a9-9fd6-478f6000b133/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1062</link>
            <description><![CDATA[<p class="para" id="N65541">ArrayExpress (https://www.ebi.ac.uk/arrayexpress) is an archive of functional genomics data at EMBL-EBI, established in 2002, initially as an archive for publication-related microarray data and was later extended to accept sequencing-based data. Over the last decade an increasing share of biological experiments involve multiple technologies assaying different biological modalities, such as epigenetics, and RNA and protein expression, and thus the BioStudies database (https://www.ebi.ac.uk/biostudies) was established to deal with such multimodal data. Its central concept is a <i>study</i>, which typically is associated with a publication. BioStudies stores metadata describing the study, provides links to the relevant databases, such as European Nucleotide Archive (ENA), as well as hosts the types of data for which specialized databases do not exist. With BioStudies now fully functional, we are able to further harmonize the archival data infrastructure at EMBL-EBI, and ArrayExpress is being migrated to BioStudies. In future, all functional genomics data will be archived at BioStudies. The process will be seamless for the users, who will continue to submit data using the online tool Annotare and will be able to query and download data largely in the same manner as before. Nevertheless, some technical aspects, particularly programmatic access, will change. This update guides the users through these changes.</p>]]></description>
            <pubDate><![CDATA[2020-11-19T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The MemMoRF database for recognizing disordered protein regions interacting with cellular membranes]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608203486-63f42805-3a43-4428-91f5-fdc3fcff2c1e/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa954</link>
            <description><![CDATA[<p class="para" id="N65541">Protein and lipid membrane interactions play fundamental roles in a large number of cellular processes (e.g. signalling, vesicle trafficking, or viral invasion). A growing number of examples indicate that such interactions can also rely on intrinsically disordered protein regions (IDRs), which can form specific reversible interactions not only with proteins but also with lipids. We named IDRs involved in such membrane lipid-induced disorder-to-order transition as MemMoRFs, in an analogy to IDRs exhibiting disorder-to-order transition upon interaction with protein partners termed Molecular Recognition Features (MoRFs). Currently, both the experimental detection and computational characterization of MemMoRFs are challenging, and information about these regions are scattered in the literature. To facilitate the related investigations we generated a comprehensive database of experimentally validated MemMoRFs based on manual curation of literature and structural data. To characterize the dynamics of MemMoRFs, secondary structure propensity and flexibility calculated from nuclear magnetic resonance chemical shifts were incorporated into the database. These data were supplemented by inclusion of sentences from papers, functional data and disease-related information. The MemMoRF database can be accessed via a user-friendly interface at https://memmorf.hegelab.org, potentially providing a central resource for the characterization of disordered regions in transmembrane and membrane-associated proteins.</p>]]></description>
            <pubDate><![CDATA[2020-10-29T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[PSORTdb 4.0: expanded and redesigned bacterial and archaeal protein subcellular localization database incorporating new secondary localizations]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608189966-4070ebf7-3b1c-4c29-a99d-976d82ab9862/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1095</link>
            <description><![CDATA[<p class="para" id="N65541">Protein subcellular localization (SCL) is important for understanding protein function, genome annotation, and aids identification of potential cell surface diagnostic markers, drug targets, or vaccine components. PSORTdb comprises ePSORTdb, a manually curated database of experimentally verified protein SCLs, and cPSORTdb, a pre-computed database of PSORTb-predicted SCLs for NCBI’s RefSeq deduced bacterial and archaeal proteomes. We now report PSORTdb 4.0 (http://db.psort.org/). It features a website refresh, in particular a more user-friendly database search. It also addresses the need to uniquely identify proteins from NCBI genomes now that GI numbers have been retired. It further expands both ePSORTdb and cPSORTdb, including additional data about novel secondary localizations, such as proteins found in bacterial outer membrane vesicles. Protein predictions in cPSORTdb have increased along with the number of available microbial genomes, from approximately 13 million when PSORTdb 3.0 was released, to over 66 million currently. Now, analyses of both complete and draft genomes are included. This expanded database will be of wide use to researchers developing SCL predictors or studying diverse microbes, including medically, agriculturally and industrially important species that have both classic or atypical cell envelope structures or vesicles.</p>]]></description>
            <pubDate><![CDATA[2020-12-11T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[CATH: increased structural coverage of functional space]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608177709-936848e0-28c4-491f-9257-253672d2442f/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1079</link>
            <description><![CDATA[<p class="para" id="N65541">CATH (https://www.cathdb.info) identifies domains in protein structures from wwPDB and classifies these into evolutionary superfamilies, thereby providing structural and functional annotations. There are two levels: CATH-B, a daily snapshot of the latest domain structures and superfamily assignments, and CATH+, with additional derived data, such as predicted sequence domains, and functionally coherent sequence subsets (Functional Families or FunFams). The latest CATH+ release, version 4.3, significantly increases coverage of structural and sequence data, with an addition of 65,351 fully-classified domains structures (+15%), providing 500 238 structural domains, and 151 million predicted sequence domains (+59%) assigned to 5481 superfamilies. The FunFam generation pipeline has been re-engineered to cope with the increased influx of data. Three times more sequences are captured in FunFams, with a concomitant increase in functional purity, information content and structural coverage. FunFam expansion increases the structural annotations provided for experimental GO terms (+59%). We also present CATH-FunVar web-pages displaying variations in protein sequences and their proximity to known or predicted functional sites. We present two case studies (1) putative cancer drivers and (2) SARS-CoV-2 proteins. Finally, we have improved links to and from CATH including SCOP, InterPro, Aquaria and 2DProt.</p>]]></description>
            <pubDate><![CDATA[2020-11-25T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[PCAT: an integrated portal for genomic and preclinical testing data of pediatric cancer patient-derived xenograft models]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608172955-63aefe37-f027-447a-ab40-1115894a16fc/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa698</link>
            <description><![CDATA[<p class="para" id="N65541">Although cancer is the leading cause of disease-related mortality in children, the relative rarity of pediatric cancers poses a significant challenge for developing novel therapeutics to further improve prognosis. Patient-derived xenograft (PDX) models, which are usually developed from high-risk tumors, are a useful platform to study molecular driver events, identify biomarkers and prioritize therapeutic agents. Here, we develop PDX for Childhood Cancer Therapeutics (PCAT), a new integrated portal for pediatric cancer PDX models. Distinct from previously reported PDX portals, PCAT is focused on pediatric cancer models and provides intuitive interfaces for querying and data mining. The current release comprises 324 models and their associated clinical and genomic data, including gene expression, mutation and copy number alteration. Importantly, PCAT curates preclinical testing results for 68 models and 79 therapeutic agents manually collected from individual agent testing studies published since 2008. To facilitate comparisons of patterns between patient tumors and PDX models, PCAT curates clinical and molecular data of patient tumors from the TARGET project. In addition, PCAT provides access to gene fusions identified in nearly 1000 TARGET samples. PCAT was built using R-shiny and MySQL. The portal can be accessed at http://pcat.zhenglab.info or http://www.pedtranscriptome.org.</p>]]></description>
            <pubDate><![CDATA[2020-08-18T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[TCRD and Pharos 2021: mining the human proteome for disease biology]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608145655-225ac362-268f-4481-96d3-ec7b48605526/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa993</link>
            <description><![CDATA[<p class="para" id="N65541">In 2014, the National Institutes of Health (NIH) initiated the Illuminating the Druggable Genome (IDG) program to identify and improve our understanding of poorly characterized proteins that can potentially be modulated using small molecules or biologics. Two resources produced from these efforts are: The Target Central Resource Database (TCRD) (http://juniper.health.unm.edu/tcrd/) and Pharos (https://pharos.nih.gov/), a web interface to browse the TCRD. The ultimate goal of these resources is to highlight and facilitate research into currently understudied proteins, by aggregating a multitude of data sources, and ranking targets based on the amount of data available, and presenting data in machine learning ready format. Since the 2017 release, both TCRD and Pharos have produced two major releases, which have incorporated or expanded an additional 25 data sources. Recently incorporated data types include human and viral-human protein–protein interactions, protein–disease and protein–phenotype associations, and drug-induced gene signatures, among others. These aggregated data have enabled us to generate new visualizations and content sections in Pharos, in order to empower users to find new areas of study in the druggable genome.</p>]]></description>
            <pubDate><![CDATA[2020-11-06T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[PANTHER version 16: a revised family classification, tree-based classification tool, enhancer regions and extensive API]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608120362-aed6e89f-b3de-41a9-be1a-b49000910c44/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1106</link>
            <description><![CDATA[<p class="para" id="N65541">PANTHER (Protein Analysis Through Evolutionary Relationships, http://www.pantherdb.org) is a resource for the evolutionary and functional classification of protein-coding genes from all domains of life. The evolutionary classification is based on a library of over 15,000 phylogenetic trees, and the functional classifications include Gene Ontology terms and pathways. Here, we analyze the current coverage of genes from genomes in different taxonomic groups, so that users can better understand what to expect when analyzing a gene list using PANTHER tools. We also describe extensive improvements to PANTHER made in the past two years. The PANTHER Protein Class ontology has been completely refactored, and 6101 PANTHER families have been manually assigned to a Protein Class, providing a high level classification of protein families and their genes. Users can access the TreeGrafter tool to add their own protein sequences to the reference phylogenetic trees in PANTHER, to infer evolutionary context as well as fine-grained annotations. We have added human enhancer-gene links that associate non-coding regions with the annotated human genes in PANTHER. We have also expanded the available services for programmatic access to PANTHER tools and data via application programming interfaces (APIs). Other improvements include additional plant genomes and an updated PANTHER GO-slim.</p>]]></description>
            <pubDate><![CDATA[2020-12-08T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[LitCovid: an open database of COVID-19 literature]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608115966-f3ba1e4b-90e6-4230-ad87-44f849698b2f/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa952</link>
            <description><![CDATA[<p class="para" id="N65541">Since the outbreak of the current pandemic in 2020, there has been a rapid growth of published articles on COVID-19 and SARS-CoV-2, with about 10 000 new articles added each month. This is causing an increasingly serious information overload, making it difficult for scientists, healthcare professionals and the general public to remain up to date on the latest SARS-CoV-2 and COVID-19 research. Hence, we developed LitCovid (https://www.ncbi.nlm.nih.gov/research/coronavirus/), a curated literature hub, to track up-to-date scientific information in PubMed. LitCovid is updated daily with newly identified relevant articles organized into curated categories. To support manual curation, advanced machine-learning and deep-learning algorithms have been developed, evaluated and integrated into the curation workflow. To the best of our knowledge, LitCovid is the first-of-its-kind COVID-19-specific literature resource, with all of its collected articles and curated data freely available. Since its release, LitCovid has been widely used, with millions of accesses by users worldwide for various information needs, such as evidence synthesis, drug discovery and text and data mining, among others.</p>]]></description>
            <pubDate><![CDATA[2020-11-09T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[GTRD: an integrated view of transcription regulation]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608091136-0472dc36-65b0-4433-ab2b-065f301c8715/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1057</link>
            <description><![CDATA[<p class="para" id="N65541">The Gene Transcription Regulation Database (GTRD; http://gtrd.biouml.org/) contains uniformly annotated and processed NGS data related to gene transcription regulation: ChIP-seq, ChIP-exo, DNase-seq, MNase-seq, ATAC-seq and RNA-seq. With the latest release, the database has reached a new level of data integration. All cell types (cell lines and tissues) presented in the GTRD were arranged into a dictionary and linked with different ontologies (BRENDA, Cell Ontology, Uberon, Cellosaurus and Experimental Factor Ontology) and with related experiments in specialized databases on transcription regulation (FANTOM5, ENCODE and GTEx). The updated version of the GTRD provides an integrated view of transcription regulation through a dedicated web interface with advanced browsing and search capabilities, an integrated genome browser, and table reports by cell types, transcription factors, and genes of interest.</p>]]></description>
            <pubDate><![CDATA[2020-11-24T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The Gene Ontology resource: enriching a GOld mine]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608056700-1c297d16-ff14-459d-a629-5d2d202bd2d3/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1113</link>
            <description><![CDATA[<p class="para" id="N65541">The Gene Ontology Consortium (GOC) provides the most comprehensive resource currently available for computable knowledge regarding the functions of genes and gene products. Here, we report the advances of the consortium over the past two years. The new GO-CAM annotation framework was notably improved, and we formalized the model with a computational schema to check and validate the rapidly increasing repository of 2838 GO-CAMs. In addition, we describe the impacts of several collaborations to refine GO and report a 10% increase in the number of GO annotations, a 25% increase in annotated gene products, and over 9,400 new scientific articles annotated. As the project matures, we continue our efforts to review older annotations in light of newer findings, and, to maintain consistency with other ontologies. As a result, 20 000 annotations derived from experimental data were reviewed, corresponding to 2.5% of experimental GO annotations. The website (http://geneontology.org) was redesigned for quick access to documentation, downloads and tools. To maintain an accurate resource and support traceability and reproducibility, we have made available a historical archive covering the past 15 years of GO data with a consistent format and file structure for both the ontology and annotations.</p>]]></description>
            <pubDate><![CDATA[2020-12-08T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[deepBase v3.0: expression atlas and interactive analysis of ncRNAs from thousands of deep-sequencing data]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608039251-65efa067-8ee2-448d-9214-4d024fd1ada4/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1039</link>
            <description><![CDATA[<p class="para" id="N65541">Eukaryotic genomes encode thousands of small and large non-coding RNAs (ncRNAs). However, the expression, functions and evolution of these ncRNAs are still largely unknown. In this study, we have updated deepBase to version 3.0 (deepBase v3.0, http://rna.sysu.edu.cn/deepbase3/index.html), an increasingly popular and openly licensed resource that facilitates integrative and interactive display and analysis of the expression, evolution, and functions of various ncRNAs by deeply mining thousands of high-throughput sequencing data from tissue, tumor and exosome samples. We updated deepBase v3.0 to provide the most comprehensive expression atlas of small RNAs and lncRNAs by integrating ∼67 620 data from 80 normal tissues and ∼50 cancer tissues. The extracellular patterns of various ncRNAs were profiled to explore their applications for discovery of noninvasive biomarkers. Moreover, we constructed survival maps of tRNA-derived RNA Fragments (tRFs), miRNAs, snoRNAs and lncRNAs by analyzing &gt;45 000 cancer sample data and corresponding clinical information. We also developed interactive webs to analyze the differential expression and biological functions of various ncRNAs in ∼50 types of cancers. This update is expected to provide a variety of new modules and graphic visualizations to facilitate analyses and explorations of the functions and mechanisms of various types of ncRNAs.</p>]]></description>
            <pubDate><![CDATA[2020-11-11T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[DualSeqDB: the host–pathogen dual RNA sequencing database for infection processes]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765603906832-7f35e860-b2ba-4af8-8d6b-b77bed4efb0e/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa890</link>
            <description><![CDATA[<p class="para" id="N65541">Despite antibiotic resistance being a matter of growing concern worldwide, the bacterial mechanisms of pathogenesis remain underexplored, restraining our ability to develop new antimicrobials. The rise of high-throughput sequencing technology has made available a massive amount of transcriptomic data that could help elucidate the mechanisms underlying bacterial infection. Here, we introduce the DualSeqDB database, a resource that helps the identification of gene transcriptional changes in both pathogenic bacteria and their natural hosts upon infection. DualSeqDB comprises nearly 300 000 entries from eight different studies, with information on bacterial and host differential gene expression under <i>in vivo</i> and <i>in vitro</i> conditions. Expression data values were calculated entirely from raw data and analyzed through a standardized pipeline to ensure consistency between different studies. It includes information on seven different strains of pathogenic bacteria and a variety of cell types and tissues in <i>Homo sapiens</i>, <i>Mus musculus</i> and <i>Macaca fascicularis</i> at different time points. We envisage that DualSeqDB can help the research community in the systematic characterization of genes involved in host infection and help the development and tailoring of new molecules against infectious diseases. DualSeqDB is freely available at http://www.tartaglialab.com/dualseq.</p>]]></description>
            <pubDate><![CDATA[2020-10-21T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[DrugSpaceX: a large screenable and synthetically tractable database extending drug space]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765603883043-39a4df70-93bd-4e18-807e-c0d9707c1629/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa920</link>
            <description><![CDATA[<p class="para" id="N65541">One of the most prominent topics in drug discovery is efficient exploration of the vast drug-like chemical space to find synthesizable and novel chemical structures with desired biological properties. To address this challenge, we created the DrugSpaceX (https://drugspacex.simm.ac.cn/) database based on expert-defined transformations of approved drug molecules. The current version of DrugSpaceX contains &gt;100 million transformed chemical products for virtual screening, with outstanding characteristics in terms of structural novelty, diversity and large three-dimensional chemical space coverage. To illustrate its practical application in drug discovery, we used a case study of discoidin domain receptor 1 (DDR1), a kinase target implicated in fibrosis and other diseases, to show DrugSpaceX performing a quick search of initial hit compounds. Additionally, for ligand identification and optimization purposes, DrugSpaceX also provides several subsets for download, including a 10% diversity subset, an extended drug-like subset, a drug-like subset, a lead-like subset, and a fragment-like subset. In addition to chemical properties and transformation instructions, DrugSpaceX can locate the position of transformation, which will enable medicinal chemists to easily integrate strategy planning and protection design.</p>]]></description>
            <pubDate><![CDATA[2020-10-26T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[CNCDatabase: a database of non-coding cancer drivers]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765603832037-872ee5e9-e971-420f-be8b-2ec34070dfe3/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa915</link>
            <description><![CDATA[<p class="para" id="N65541">Most mutations in cancer genomes occur in the non-coding regions with unknown impact on tumor development. Although the increase in the number of cancer whole-genome sequences has revealed numerous putative non-coding cancer drivers, their information is dispersed across multiple studies making it difficult to understand their roles in tumorigenesis of different cancer types. We have developed CNCDatabase, Cornell Non-coding Cancer driver Database (https://cncdatabase.med.cornell.edu/) that contains detailed information about predicted non-coding drivers at gene promoters, 5′ and 3′ UTRs (untranslated regions), enhancers, CTCF insulators and non-coding RNAs. CNCDatabase documents 1111 protein-coding genes and 90 non-coding RNAs with reported drivers in their non-coding regions from 32 cancer types by computational predictions of positive selection using whole-genome sequences; differential gene expression in samples with and without mutations; or another set of experimental validations including luciferase reporter assays and genome editing. The database can be easily modified and scaled as lists of non-coding drivers are revised in the community with larger whole-genome sequencing studies, CRISPR screens and further experimental validations. Overall, CNCDatabase provides a helpful resource for researchers to explore the pathological role of non-coding alterations in human cancers.</p>]]></description>
            <pubDate><![CDATA[2020-10-23T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[UniProt: the universal protein knowledgebase in 2021]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765603814441-494a9f23-5ed5-42cf-8c84-6a544007b00c/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1100</link>
            <description><![CDATA[<p class="para" id="N65541">The aim of the UniProt Knowledgebase is to provide users with a comprehensive, high-quality and freely accessible set of protein sequences annotated with functional information. In this article, we describe significant updates that we have made over the last two years to the resource. The number of sequences in UniProtKB has risen to approximately 190 million, despite continued work to reduce sequence redundancy at the proteome level. We have adopted new methods of assessing proteome completeness and quality. We continue to extract detailed annotations from the literature to add to reviewed entries and supplement these in unreviewed entries with annotations provided by automated systems such as the newly implemented Association-Rule-Based Annotator (ARBA). We have developed a credit-based publication submission interface to allow the community to contribute publications and annotations to UniProt entries. We describe how UniProtKB responded to the COVID-19 pandemic through expert curation of relevant entries that were rapidly made available to the research community through a dedicated portal. UniProt resources are available under a CC-BY (4.0) license via the web at https://www.uniprot.org/.</p>]]></description>
            <pubDate><![CDATA[2020-11-25T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[3DIV update for 2021: a comprehensive resource of 3D genome and 3D cancer genome]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765602167220-7f5fc953-18c9-4ca6-b138-e4428935b798/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1078</link>
            <description><![CDATA[<p class="para" id="N65541">Three-dimensional (3D) genome organization is tightly coupled with gene regulation in various biological processes and diseases. In cancer, various types of large-scale genomic rearrangements can disrupt the 3D genome, leading to oncogenic gene expression. However, unraveling the pathogenicity of the 3D cancer genome remains a challenge since closer examinations have been greatly limited due to the lack of appropriate tools specialized for disorganized higher-order chromatin structure. Here, we updated a 3D-genome Interaction Viewer and database named 3DIV by uniformly processing ∼230 billion raw Hi-C reads to expand our contents to the 3D cancer genome. The updates of 3DIV are listed as follows: (i) the collection of 401 samples including 220 cancer cell line/tumor Hi-C data, 153 normal cell line/tissue Hi-C data, and 28 promoter capture Hi-C data, (ii) the live interactive manipulation of the 3D cancer genome to simulate the impact of structural variations and (iii) the reconstruction of Hi-C contact maps by user-defined chromosome order to investigate the 3D genome of the complex genomic rearrangement. In summary, the updated 3DIV will be the most comprehensive resource to explore the gene regulatory effects of both the normal and cancer 3D genome. ‘3DIV’ is freely available at http://3div.kr.</p>]]></description>
            <pubDate><![CDATA[2020-11-27T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[RepeatsDB in 2021: improved data and extended classification for protein tandem repeat structures]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765602155498-48010892-811d-4565-8ed1-117c8c602234/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1097</link>
            <description><![CDATA[<p class="para" id="N65541">The RepeatsDB database (URL: https://repeatsdb.org/) provides annotations and classification for protein tandem repeat structures from the Protein Data Bank (PDB). Protein tandem repeats are ubiquitous in all branches of the tree of life. The accumulation of solved repeat structures provides new possibilities for classification and detection, but also increasing the need for annotation. Here we present RepeatsDB 3.0, which addresses these challenges and presents an extended classification scheme. The major conceptual change compared to the previous version is the hierarchical classification combining top levels based solely on structural similarity (Class &gt; Topology &gt; Fold) with two new levels (Clan &gt; Family) requiring sequence similarity and describing repeat motifs in collaboration with Pfam. Data growth has been addressed with improved mechanisms for browsing the classification hierarchy. A new UniProt-centric view unifies the increasingly frequent annotation of structures from identical or similar sequences. This update of RepeatsDB aligns with our commitment to develop a resource that extracts, organizes and distributes specialized information on tandem repeat protein structures.</p>]]></description>
            <pubDate><![CDATA[2020-11-25T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[iModulonDB: a knowledgebase of microbial transcriptional regulation derived from machine learning]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765602151124-cb4bd30f-1796-4613-86ce-9047e4b5c9e8/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa810</link>
            <description><![CDATA[<p class="para" id="N65541">Independent component analysis (ICA) of bacterial transcriptomes has emerged as a powerful tool for obtaining co-regulated, independently-modulated gene sets (iModulons), inferring their activities across a range of conditions, and enabling their association to known genetic regulators. By grouping and analyzing genes based on observations from big data alone, iModulons can provide a novel perspective into how the composition of the transcriptome adapts to environmental conditions. Here, we present iModulonDB (imodulondb.org), a knowledgebase of prokaryotic transcriptional regulation computed from high-quality transcriptomic datasets using ICA. Users select an organism from the home page and then search or browse the curated iModulons that make up its transcriptome. Each iModulon and gene has its own interactive dashboard, featuring plots and tables with clickable, hoverable, and downloadable features. This site enhances research by presenting scientists of all backgrounds with co-expressed gene sets and their activity levels, which lead to improved understanding of regulator-gene relationships, discovery of transcription factors, and the elucidation of unexpected relationships between conditions and genetic regulatory activity. The current release of iModulonDB covers three organisms (<i>Escherichia coli, Staphylococcus aureus</i> and <i>Bacillus subtilis</i>) with 204 iModulons, and can be expanded to cover many additional organisms.</p>]]></description>
            <pubDate><![CDATA[2020-10-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[GenBank]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765602148071-a8d1db33-9674-457b-898a-45a24cec6a0e/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1023</link>
            <description><![CDATA[<p class="para" id="N65541">GenBank<sup>®</sup> (https://www.ncbi.nlm.nih.gov/genbank/) is a comprehensive, public database that contains 9.9 trillion base pairs from over 2.1 billion nucleotide sequences for 478 000 formally described species. Daily data exchange with the European Nucleotide Archive and the DNA Data Bank of Japan ensures worldwide coverage. Recent updates include new resources for data from the SARS-CoV-2 virus, updates to the NCBI Submission Portal and associated submission wizards for dengue and SARS-CoV-2 viruses, new taxonomy queries for viruses and prokaryotes, and simplified submission processes for EST and GSS sequences.</p>]]></description>
            <pubDate><![CDATA[2020-11-16T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[DIANA-miRGen v4: indexing promoters and regulators for more than 1500 microRNAs]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765602143791-e4900677-7def-452b-a176-bfe4b4d27dfb/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1060</link>
            <description><![CDATA[<p class="para" id="N65541">Deregulation of microRNA (miRNA) expression plays a critical role in the transition from a physiological to a pathological state. The accurate miRNA promoter identification in multiple cell types is a fundamental endeavor towards understanding and characterizing the underlying mechanisms of both physiological as well as pathological conditions. DIANA-miRGen v4 (www.microrna.gr/mirgenv4) provides cell type specific miRNA transcription start sites (TSSs) for over 1500 miRNAs retrieved from the analysis of &gt;1000 cap analysis of gene expression (CAGE) samples corresponding to 133 tissues, cell lines and primary cells available in FANTOM repository. MiRNA TSS locations were associated with transcription factor binding site (TFBSs) annotation, for &gt;280 TFs, derived from analyzing the majority of ENCODE ChIP-Seq datasets. For the first time, clusters of cell types having common miRNA TSSs are characterized and provided through a user friendly interface with multiple layers of customization. DIANA-miRGen v4 significantly improves our understanding of miRNA biogenesis regulation at the transcriptional level by providing a unique integration of high-quality annotations for hundreds of cell specific miRNA promoters with experimentally derived TFBSs.</p>]]></description>
            <pubDate><![CDATA[2020-11-27T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[NASA GeneLab: interfaces for the exploration of space omics data]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765602130993-caac6cb1-3ae3-4413-b12d-30f9c56574fa/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa887</link>
            <description><![CDATA[<p class="para" id="N65541">The mission of NASA’s GeneLab database (https://genelab.nasa.gov/) is to collect, curate, and provide access to the genomic, transcriptomic, proteomic and metabolomic (so-called ‘omics’) data from biospecimens flown in space or exposed to simulated space stressors, maximizing their utilization. This large collection of data enables the exploration of molecular network responses to space environments using a systems biology approach. We review here the various components of the GeneLab platform, including the new data repository web interface, and the GeneLab Online Data Entry (GEODE) web portal, which will support the expansion of the database in the future to include companion non-omics assay data. We discuss our design for GEODE, particularly how it promotes investigators providing more accurate metadata, reducing the curation effort required of GeneLab staff. We also introduce here a new GeneLab Application Programming Interface (API) specifically designed to support tools for the visualization of processed omics data. We review the outreach efforts by GeneLab to utilize the spaceflight data in the repository to generate novel discoveries and develop new hypotheses, including spearheading data analysis working groups, and a high school student training program. All these efforts are aimed ultimately at supporting precision risk management for human space exploration.</p>]]></description>
            <pubDate><![CDATA[2020-10-20T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The Zebrafish Information Network: major gene page and home page updates]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765602084677-043a856b-27ac-4bc6-9e04-0ffda80ba341/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa1010</link>
            <description><![CDATA[<p class="para" id="N65541">The Zebrafish Information Network (ZFIN) (https://zfin.org/) is the database for the model organism, zebrafish (<i>Danio rerio</i>). ZFIN expertly curates, organizes, and provides a wide array of zebrafish genetic and genomic data, including genes, alleles, transgenic lines, gene expression, gene function, mutant phenotypes, orthology, human disease models, gene and mutant nomenclature, and reagents. New features at ZFIN include major updates to the home page and the gene page, the two most used pages at ZFIN. Data including disease models, phenotypes, expression, mutants and gene function continue to be contributed to The Alliance of Genome Resources for integration with similar data from other model organisms.</p>]]></description>
            <pubDate><![CDATA[2020-11-10T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The InterPro protein families and domains database: 20 years on]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765602077586-ba26a897-b043-4785-b315-52de814e48d4/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa977</link>
            <description><![CDATA[<p class="para" id="N65541">The InterPro database (https://www.ebi.ac.uk/interpro/) provides an integrative classification of protein sequences into families, and identifies functionally important domains and conserved sites. InterProScan is the underlying software that allows protein and nucleic acid sequences to be searched against InterPro's signatures. Signatures are predictive models which describe protein families, domains or sites, and are provided by multiple databases. InterPro combines signatures representing equivalent families, domains or sites, and provides additional information such as descriptions, literature references and Gene Ontology (GO) terms, to produce a comprehensive resource for protein classification. Founded in 1999, InterPro has become one of the most widely used resources for protein family annotation. Here, we report the status of InterPro (version 81.0) in its 20th year of operation, and its associated software, including updates to database content, the release of a new website and REST API, and performance improvements in InterProScan.</p>]]></description>
            <pubDate><![CDATA[2020-11-06T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[SC2disease: a manually curated database of single-cell transcriptome for human diseases]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765602052758-b2f8b9ef-d6f5-4464-9188-a37085235d91/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1093/nar/gkaa838</link>
            <description><![CDATA[<p class="para" id="N65541">SC2disease (http://easybioai.com/sc2disease/) is a manually curated database that aims to provide a comprehensive and accurate resource of gene expression profiles in various cell types for different diseases. With the development of single-cell RNA sequencing (scRNA-seq) technologies, uncovering cellular heterogeneity of different tissues for different diseases has become feasible by profiling transcriptomes across cell types at the cellular level. In particular, comparing gene expression profiles between different cell types and identifying cell-type-specific genes in various diseases offers new possibilities to address biological and medical questions. However, systematic, hierarchical and vast databases of gene expression profiles in human diseases at the cellular level are lacking. Thus, we reviewed the literature prior to March 2020 for studies which used scRNA-seq to study diseases with human samples, and developed the SC2disease database to summarize all the data by different diseases, tissues and cell types. SC2disease documents 946 481 entries, corresponding to 341 cell types, 29 tissues and 25 diseases. Each entry in the SC2disease database contains comparisons of differentially expressed genes between different cell types, tissues and disease-related health status. Furthermore, we reanalyzed gene expression matrix by unified pipeline to improve the comparability between different studies. For each disease, we also compare cell-type-specific genes with the corresponding genes of lead single nucleotide polymorphisms (SNPs) identified in genome-wide association studies (GWAS) to implicate cell type specificity of the traits.</p>]]></description>
            <pubDate><![CDATA[2020-10-03T00:00]]></pubDate>
        </item>
    </channel>
</rss>