<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:media="http://search.yahoo.com/mrss/" xmlns:ynews="http://news.yahoo.com/rss/">
    <channel>
        <title>Nova Reader - Subject</title>
        <link>https://www.novareader.co</link>
        <description>Default RSS Feed</description>
        <language>en-us</language>
        <copyright>Newgen KnowledgeWorks</copyright>
        <item>
            <title><![CDATA[The implementation of random survival forests in conflict management data: An examination of power sharing and third party mediation in post-conflict countries]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766072629226-474375f5-fd22-4325-80af-a421be3b910c/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0250963</link>
            <description><![CDATA[<p class="para" id="N65539">Time-to-event analysis is a common occurrence in political science. In recent years, there has been an increased usage of machine learning methods in quantitative political science research. This article advocates for the implementation of machine learning duration models to assist in a sound model selection process. We provide a brief tutorial introduction to the random survival forest (RSF) algorithm and contrast it to a popular predecessor, the Cox proportional hazards model, with emphasis on methodological utility for political science researchers. We implement both methods for simulated time-to-event data and the Power-Sharing Event Dataset (PSED) to assist researchers in evaluating the merits of machine learning duration models. We provide evidence of significantly higher survival probabilities for peace agreements with 3rd party mediated design and implementation. We also detect increased survival probabilities for peace agreements that incorporate territorial power-sharing and avoid multiple rebel party signatories. Further, the RSF, a previously under-used method for analyzing political science time-to event data, provides a novel approach for ranking of peace agreement criteria importance in predicting peace agreement duration. Our findings demonstrate a scenario exhibiting the interpretability and performance of RSF for political science time-to-event data. These findings justify the robust interpretability and competitive performance of the random survival forest algorithm in numerous circumstances, in addition to promoting a diverse, holistic model-selection process for time-to-event political science data.</p>]]></description>
            <pubDate><![CDATA[2021-05-03T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Application of machine learning and genetic optimization algorithms for modeling and optimizing soybean yield using its component traits]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766072600516-c09168fb-5f68-4d29-9423-5a3fd4471c5b/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0250665</link>
            <description><![CDATA[<p class="para" id="N65539">Improving genetic yield potential in major food grade crops such as soybean <i>(Glycine max</i> L.) is the most sustainable way to address the growing global food demand and its security concerns. Yield is a complex trait and reliant on various related variables called yield components. In this study, the five most important yield component traits in soybean were measured using a panel of 250 genotypes grown in four environments. These traits were the number of nodes per plant (NP), number of non-reproductive nodes per plant (NRNP), number of reproductive nodes per plant (RNP), number of pods per plant (PP), and the ratio of number of pods to number of nodes per plant (P/N). These data were used for predicting the total soybean seed yield using the Multilayer Perceptron (MLP), Radial Basis Function (RBF), and Random Forest (RF), machine learning (ML) algorithms, individually and collectively through an ensemble method based on bagging strategy (E-B). The RBF algorithm with highest Coefficient of Determination (R<sup>2</sup>) value of 0.81 and the lowest Mean Absolute Errors (MAE) and Root Mean Square Error (RMSE) values of 148.61 kg.ha<sup>-1</sup>, and 185.31 kg.ha<sup>-1</sup>, respectively, was the most accurate algorithm and, therefore, selected as the metaClassifier for the E-B algorithm. Using the E-B algorithm, we were able to increase the prediction accuracy by improving the values of R<sup>2</sup>, MAE, and RMSE by 0.1, 0.24 kg.ha<sup>-1</sup>, and 0.96 kg.ha<sup>-1</sup>, respectively. Furthermore, for the first time in this study, we allied the E-B with the genetic algorithm (GA) to model the optimum values of yield components in an ideotype genotype in which the yield is maximized. The results revealed a better understanding of the relationships between soybean yield and its components, which can be used for selecting parental lines and designing promising crosses for developing cultivars with improved genetic yield potential.</p>]]></description>
            <pubDate><![CDATA[2021-04-30T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Convolutional neural networks improve species distribution modelling by capturing the spatial structure of the environment]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766072007211-fc03c380-be94-44a7-bd10-a6b45c3954c4/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pcbi.1008856</link>
            <description><![CDATA[<p class="para" id="N65539">Convolutional Neural Networks (CNNs) are statistical models suited for learning complex visual patterns. In the context of Species Distribution Models (SDM) and in line with predictions of landscape ecology and island biogeography, CNN could grasp how local landscape structure affects prediction of species occurrence in SDMs. The prediction can thus reflect the signatures of entangled ecological processes. Although previous machine-learning based SDMs can learn complex influences of environmental predictors, they cannot acknowledge the influence of environmental structure in local landscapes (hence denoted “punctual models”). In this study, we applied CNNs to a large dataset of plant occurrences in France (GBIF), on a large taxonomical scale, to predict ranked relative probability of species (by joint learning) to any geographical position. We examined the way local environmental landscapes improve prediction by performing alternative CNN models deprived of information on landscape heterogeneity and structure (“ablation experiments”). We found that the landscape structure around location crucially contributed to improve predictive performance of CNN-SDMs. CNN models can classify the predicted distributions of many species, as other joint modelling approaches, but they further prove efficient in identifying the influence of local environmental landscapes. CNN can then represent signatures of spatially structured environmental drivers. The prediction gain is noticeable for rare species, which open promising perspectives for biodiversity monitoring and conservation strategies. Therefore, the approach is of both theoretical and practical interest. We discuss the way to test hypotheses on the patterns learnt by CNN, which should be essential for further interpretation of the ecological processes at play.</p><p class="para" id="N65542">Species distribution models aim at linking species spatial distribution to the environment. They can highlight the ecological preferences of species and thus predict which species are likely to be present in a given environment. These models are used in many scenarios such as conservation plans or monitoring of invasive species. The choice of model and the environmental data used have a strong impact on the model’s ability to capture important information. Specificaly, state-of-the-art models generally use a punctual environment and do not take into account the environmental context or neighbourhood. Here we present a species distribution model based on a convolutional neural network that allows the use of large scale data such as spatialized environmental data including the environmental neighbourhood in addition to the punctual environment. We highlight the interests and limitations of this method as well as the importance of the environmental context in learning about species distributions.</p>]]></description>
            <pubDate><![CDATA[2021-04-19T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Spectrum decomposition in Gaussian scale space for uneven illumination image binarization]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766071625882-555f0b3b-881d-4b6f-9ab7-8664a71a45fa/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0251014</link>
            <description><![CDATA[<p class="para" id="N65539">Although most images in industrial applications have fewer targets and simple image backgrounds, binarization is still a challenging task, and the corresponding results are usually unsatisfactory because of uneven illumination interference. In order to efficiently threshold images with nonuniform illumination, this paper proposes an efficient global binarization algorithm that estimates the inhomogeneous background surface of the original image constructed from the first <i>k</i> leading principal components in the Gaussian scale space (GSS). Then, we use the difference operator to extract the distinct foreground of the original image in which the interference of uneven illumination is effectively eliminated. Finally, the image can be effortlessly binarized by an existing global thresholding algorithm such as the Otsu method. In order to qualitatively and quantitatively verify the segmentation performance of the presented scheme, experiments were performed on a dataset collected from a nonuniform illumination environment. Compared with classical binarization methods, in some metrics, the experimental results demonstrate the effectiveness of the introduced algorithm in providing promising binarization outcomes and low computational costs.</p>]]></description>
            <pubDate><![CDATA[2021-04-30T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Predicting the animal hosts of coronaviruses from compositional biases of spike protein and whole genome sequences through machine learning]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766071191315-5d26c526-0d4e-4778-86da-863b40a5ae4b/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.ppat.1009149</link>
            <description><![CDATA[<p class="para" id="N65539">The COVID-19 pandemic has demonstrated the serious potential for novel zoonotic coronaviruses to emerge and cause major outbreaks. The immediate animal origin of the causative virus, SARS-CoV-2, remains unknown, a notoriously challenging task for emerging disease investigations. Coevolution with hosts leads to specific evolutionary signatures within viral genomes that can inform likely animal origins. We obtained a set of 650 spike protein and 511 whole genome nucleotide sequences from 222 and 185 viruses belonging to the family <i>Coronaviridae</i>, respectively. We then trained random forest models independently on genome composition biases of spike protein and whole genome sequences, including dinucleotide and codon usage biases in order to predict animal host (of nine possible categories, including human). In hold-one-out cross-validation, predictive accuracy on unseen coronaviruses consistently reached ~73%, indicating evolutionary signal in spike proteins to be just as informative as whole genome sequences. However, different composition biases were informative in each case. Applying optimised random forest models to classify human sequences of MERS-CoV and SARS-CoV revealed evolutionary signatures consistent with their recognised intermediate hosts (camelids, carnivores), while human sequences of SARS-CoV-2 were predicted as having bat hosts (suborder Yinpterochiroptera), supporting bats as the suspected origins of the current pandemic. In addition to phylogeny, variation in genome composition can act as an informative approach to predict emerging virus traits as soon as sequences are available. More widely, this work demonstrates the potential in combining genetic resources with machine learning algorithms to address long-standing challenges in emerging infectious diseases.</p><p class="para" id="N65542">New zoonotic viruses remain a major threat to global health and the COVID-19 pandemic has shown the specific potential of coronaviruses to cause widespread disease burden and economic damage. Tracing the origins of these zoonotic viruses is extremely challenging and usually requires substantial effort. However, there is potential to uncover which animals may be the host origin of viruses by using ‘signatures’ within viral genomes generated by long-term coevolution. We investigated this by calculating 116 genomic features of spike protein sequences and whole genome sequences from approximately 200 coronaviruses. We used a machine learning approach in random forests, training separate models to predict broad host type using genomic information from spike proteins or whole genomes. Models trained on spike proteins achieved similar performance to that of whole genomes, reiterating the importance of this protein for host-virus interactions and likelihood of cross-species transmission. When applied to SARS-CoV-2, the causative virus of COVID-19, model predictions suggested a bat origin, consistent with estimations elsewhere using more traditional phylogenetic analyses. This work demonstrates the potential of machine learning to infer the ecology of new zoonotic viruses directly from genetic sequences, giving a rapid methodology to assist in tracing the origins of outbreaks.</p>]]></description>
            <pubDate><![CDATA[2021-04-20T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Prediction of the compressive strength of high-performance self-compacting concrete by an ultrasonic-rebound method based on a GA-BP neural network]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766071097674-5a268c94-933d-410e-9691-dd792a560246/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0250795</link>
            <description><![CDATA[<p class="para" id="N65539">To address the problem of low accuracy and poor robustness of in situ testing of the compressive strength of high-performance self-compacting concrete (SCC), a genetic algorithm (GA)-optimized backpropagation neural network (BPNN) model was established to predict the compressive strength of SCC. Experiments based on two concrete nondestructive testing methods, i.e., ultrasonic pulse velocity and Schmidt rebound hammer, were designed and test sample data were obtained. A neural network topology with two input nodes, 19 hidden nodes, and one output node was constructed, and the initial weights and thresholds of the resulting traditional BPNN model were optimized using GA. The results showed a correlation coefficient of 0.967 between the values predicted by the established BPNN model and the test values, with an RMSE of 3.703, compared to a correlation coefficient of 0.979 between the values predicted by the GA-optimized BPNN model and the test values, with an RMSE of 2.972. The excellent agreement between the predicted and test values demonstrates the model can accurately predict the compressive strength of SCC and hence reduce the cost and time for SCC compressive strength testing.</p>]]></description>
            <pubDate><![CDATA[2021-05-03T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[A decentralized hybrid computing consumer authentication framework for a reliable drone delivery as a service]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766070989878-01cfb966-0de3-4f29-b93a-97da70a9f511/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0250737</link>
            <description><![CDATA[<p class="para" id="N65539">The thriving adoption of drones for delivering parcels, packages, medicines, etc., is surging with time. The application of drones for delivery services results in faster delivery, fuel-saving, and less energy consumption. Giant companies like Google, Amazon, Facebook, etc., are actively working on developing, testing, and improving drone-based delivery systems. So far, a lot of work has been done for improving the design, speed, operating range, security of the delivery drones, etc. However, very limited work has been done to ensure a complete and reliable last-mile delivery from the merchant’s store to the hands of the actual customer. To ensure a complete and reliable last-mile delivery, a drone must authenticate the consumer before dropping the package. Therefore, in this work, we propose a consumer authentication (Consumer-Auth) hybrid computing framework for drone delivery as a service to make sure that the parcel is perfectly delivered to the intended customer. The proposed Consumer-Auth framework enables a drone to reach the exact destination by using the GPS coordinates of the customer autonomously. After reaching the exact location, the drone waits for the customer to come to the specific pinned location then it starts a two-factor consumer authentication process, i.e., one-time password (OTP) verification and face Recognition. The experimental results manifest the effectiveness of the proposed Consumer-Auth framework to ensure a complete and reliable drone-based last-mile delivery.</p>]]></description>
            <pubDate><![CDATA[2021-04-30T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[A computational method for drug sensitivity prediction of cancer cell lines based on various molecular information]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766069976981-f60b4797-f320-4efa-b6d1-8b7ddea16380/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0250620</link>
            <description><![CDATA[<p class="para" id="N65539">Determining sensitive drugs for a patient is one of the most critical problems in precision medicine. Using genomic profiles of the tumor and drug information can help in tailoring the most efficient treatment for a patient. In this paper, we proposed a classification machine learning approach that predicts the sensitive/resistant drugs for a cell line. It can be performed by using both drug and cell line similarities, one of the cell line or drug similarities, or even not using any similarity information. This paper investigates the influence of using previously defined as well as two newly introduced similarities on predicting anti-cancer drug sensitivity. The proposed method uses max concentration thresholds for assigning drug responses to class labels. Its performance was evaluated using stratified five-fold cross-validation on cell line-drug pairs in two datasets. Assessing the predictive powers of the proposed model and three sets of methods, including state-of-the-art classification methods, state-of-the-art regression methods, and off-the-shelf classification machine learning approaches shows that the proposed method outperforms other methods. Moreover, The efficiency of the model is evaluated in tissue-specific conditions. Besides, the novel sensitive associations predicted by this model were verified by several supportive evidence in the literature and reliable database. Therefore, the proposed model can efficiently be used in predicting anti-cancer drug sensitivity. Material and implementation are available at https://github.com/fahmadimoughari/CDSML.</p>]]></description>
            <pubDate><![CDATA[2021-04-29T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Predicting corporate credit risk: Network contagion via trade credit]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766069909408-4294bb47-e1b9-493d-8100-f728dc417820/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0250115</link>
            <description><![CDATA[<p class="para" id="N65539">Trade credit is a payment extension granted by a selling firm to its customer. Companies typically respond to late payments from their customers by delaying payments to suppliers, thus generating a ripple through the transaction network. Therefore, trade credit is as a potential vehicle of propagation of losses in case of default events. The goal of this work is to leverage information on the trade credit among connected firms to predict imminent defaults of firms. We use a unique dataset of client firms of a major Italian bank to investigate firm bankruptcy between October 2016 to March 2018. We develop a model to capture network spillover effects originating from the supply chain on the probability of default of each firm via a sequential approach: the output of a first model component on single firm features is used in a subsequent model which captures network spillovers. While the first component is the standard econometrics way to predict such dynamics, the network module represents an innovative way to look into the effect of trade credit on default probability. This module looks at the transaction network of the firm, as inferred from the payments transiting via the bank, in order to identify the trade partners of the firm. By using several features extracted from the network of transactions, this model is able to predict a large fraction of the defaults, thus showing the value hidden in the network information. Finally, we merge firm and network features with a machine learning model to create a ‘hybrid’ model, which improves the recall for the task by almost 20 percentage points over the baseline.</p>]]></description>
            <pubDate><![CDATA[2021-04-29T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Gastric polyp detection in gastroscopic images using deep neural network]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766069897414-ef474342-0f45-4ee4-a057-b276273b559e/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0250632</link>
            <description><![CDATA[<p class="para" id="N65539">This paper presents the research results of detecting gastric polyps with deep learning object detection method in gastroscopic images. Gastric polyps have various sizes. The difficulty of polyp detection is that small polyps are difficult to detect from the background. We propose a feature extraction and fusion module and combine it with the YOLOv3 network to form our network. This method performs better than other methods in the detection of small polyps because it can fuse the semantic information of high-level feature maps with low-level feature maps to help small polyps detection. In this work, we use a dataset of gastric polyps created by ourselves, containing 1433 training images and 508 validation images. We train and validate our network on our dataset. In comparison with other methods of polyps detection, our method has a significant improvement in precision, recall rate, F1, and F2 score. The precision, recall rate, F1 score, and F2 score of our method can achieve 91.6%, 86.2%, 88.8%, and 87.2%.</p>]]></description>
            <pubDate><![CDATA[2021-04-28T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Random forest model for feature-based Alzheimer’s disease conversion prediction from early mild cognitive impairment subjects]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766069550536-d9a4c7df-3040-4ee9-bf7b-2d09a087cea5/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244773</link>
            <description><![CDATA[<p class="para" id="N65539">Alzheimer’s Disease (AD) conversion prediction from the mild cognitive impairment (MCI) stage has been a difficult challenge. This study focuses on providing an individualized MCI to AD conversion prediction using a balanced random forest model that leverages clinical data. In order to do this, 383 Early Mild Cognitive Impairment (EMCI) patients were gathered from the Alzheimer’s Disease Neuroimaging Initiative (ADNI). Of these patients, 49 would eventually convert to AD (EMCI_C), whereas the remaining 334 did not convert (EMCI_NC). All of these patients were split randomly into training and testing data sets with 95 patients reserved for testing. Nine clinical features were selected, comprised of a mix of demographic, brain volume, and cognitive testing variables. Oversampling was then performed in order to balance the initially imbalanced classes prior to training the model with 1000 estimators. Our results showed that a random forest model was effective (93.6% accuracy) at predicting the conversion of EMCI patients to AD based on these clinical features. Additionally, we focus on explainability by assessing the importance of each clinical feature. Our model could impact the clinical environment as a tool to predict the conversion to AD from a prodromal stage or to identify ideal candidates for clinical trials.</p>]]></description>
            <pubDate><![CDATA[2021-04-29T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Impact of pulsed-wave-Doppler velocity-envelope tracing techniques on classification of complete fetal cardiac cycles]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766069281965-86084a69-99f2-4f6e-be7f-9a94ac01d3b6/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0248114</link>
            <description><![CDATA[<p class="para" id="N65539">Fetal echocardiography is an operator-dependent examination technique requiring a high level of expertise. Pulsed-wave Doppler (PWD) is often used as a reference for the mechanical activity of the heart, from which several quantitative parameters can be extracted. These aspects suggest the development of software tools that can reliably identify complete and clinically meaningful fetal cardiac cycles that can enable their automatic measurement. Several scientific works have addressed the tracing of the PWD velocity envelope. In this work, we assess the different steps involved in the signal processing chains that enable PWD envelope tracing. We apply a supervised classifier trained on envelopes traced by different signal processing chains for distinguishing complete and measurable PWD heartbeats from incomplete or malformed ones, which makes it possible to determine the impact of each of the different processing steps on the detection accuracy. In this study, we collected 43 images and labeled 174,319 PWD segments from 25 pregnant women volunteers. By considering seven envelope tracing techniques and the 23 different processing steps involved in their implementation, the results of our study reveal that, compared to the steps investigated in most other works, those that achieve binarisation and envelope extraction are significantly more important (<i>p</i> &lt; 0.05). The best approaches among those studied enabled greater than 98% accuracy on our large manually annotated dataset.</p>]]></description>
            <pubDate><![CDATA[2021-04-28T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Effects of subthreshold nanosecond laser therapy in age-related macular degeneration using artificial intelligence (STAR-AI Study)]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766069153775-670a1f8d-338c-47a1-a67a-6aef4ff36a02/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0250609</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Purpose</h3><p class="para" id="N65543">To investigate changes in retinal thickness, drusen volume, and visual acuity following subthreshold nanosecond laser (SNL) treatment in patients with age-related macular degeneration (ARMD).</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Design</h3><p class="para" id="N65549">Retrospective chart review.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Methods</h3><p class="para" id="N65555">Patients with intermediate ARMD treated with a single session of SNL (2RT®, Ellex R&amp;D Pty Ltd, Adelaide, Australia) were included. Swept-source optical coherence tomography (OCT) imaging (Triton; Topcon Medical Systems, Tokyo, Japan) was performed within 6 months before and after SNL treatment. Retinal layers were segmented using the artificial intelligence-enabled Orion® software (Voxeleron LLC, San Francisco, USA). The macular region was analyzed according to the Early Treatment Diabetic Retinopathy Study map. Mean difference and standard deviation in baseline and post-treatment retinal layer thicknesses are reported.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Results</h3><p class="para" id="N65561">37 eyes from 25 patients were included in this study (mean age 74.7±9.2 years). An average of 51±6 spots were applied around the macula of each study eye, with a mean spot power of 0.33±0.04mJ. Increases in total retinal thickness were observed within the outer temporal and inferior sectors (P&lt;0.05). Within the annulus, there was an increase in thickness of the sub-retinal pigment epithelial (RPE) space [0.88±2.41μm, P = 0.03], defined between the RPE and Bruch’s membrane. An increase in thickness of 1.13±2.55μm (P = 0.01) was also noted in the inferior sector of the photoreceptor complex, defined from the inner and outer segment junction to the RPE. Decreases in thickness were observed within the superior sector of the inner nuclear layer (INL) [-1.08±2.55μm, P = 0.01], and within the annulus of the outer nuclear layer (ONL) [-1.44±3.55μm, P = 0.02].</p></div><div class="section" id="sec005"><h3 class="BHead" id="nov000-5">Conclusions</h3><p class="para" id="N65567">At 6 months post-SNL treatment, there were sectoral increases in OPL, photoreceptor complex, and sub-RPE space thicknesses and sectoral decreases in INL and ONL thicknesses. This pilot study demonstrates the utility of OCT combined with artificial intelligence-enabled software to track retinal changes that occur following SNL treatment in intermediate ARMD.</p></div>]]></description>
            <pubDate><![CDATA[2021-04-29T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Accurate cancer phenotype prediction with AKLIMATE, a stacked kernel learner integrating multimodal genomic data and pathway knowledge]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766068822243-08230175-e063-42d1-9e28-81ef17d4ff9c/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pcbi.1008878</link>
            <description><![CDATA[<p class="para" id="N65539">Advancements in sequencing have led to the proliferation of multi-omic profiles of human cells under different conditions and perturbations. In addition, many databases have amassed information about pathways and gene “signatures”—patterns of gene expression associated with specific cellular and phenotypic contexts. An important current challenge in systems biology is to leverage such knowledge about gene coordination to maximize the predictive power and generalization of models applied to high-throughput datasets. However, few such integrative approaches exist that also provide interpretable results quantifying the importance of individual genes and pathways to model accuracy. We introduce AKLIMATE, a first kernel-based stacked learner that seamlessly incorporates multi-omics feature data with prior information in the form of pathways for either regression or classification tasks. AKLIMATE uses a novel multiple-kernel learning framework where individual kernels capture the prediction propensities recorded in random forests, each built from a specific pathway gene set that integrates all omics data for its member genes. AKLIMATE has comparable or improved performance relative to state-of-the-art methods on diverse phenotype learning tasks, including predicting microsatellite instability in endometrial and colorectal cancer, survival in breast cancer, and cell line response to gene knockdowns. We show how AKLIMATE is able to connect feature data across data platforms through their common pathways to identify examples of several known and novel contributors of cancer and synthetic lethality.</p><p class="para" id="N65542">We describe a new method that incorporates multimodal molecular measurements with gene pathway information for classification and regression prediction tasks. The method combines powerful machine learning methodologies such as ensemble, kernel and stacked learning. A key new contribution is the creation of empirical kernels from pairwise similarities of sample predictions and sample paths along the trees of a random forest specific to a particular gene module. We demonstrate the method performs as well as, or better than, top approaches on three very different cancer genomics applications.</p>]]></description>
            <pubDate><![CDATA[2021-04-16T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Validity concerns with the Revised Study Process Questionnaire (R-SPQ-2F) in undergraduate anatomy &amp; physiology students]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766068588249-008db868-7f82-4fb5-ba50-f469e896c6b5/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0250600</link>
            <description><![CDATA[<p class="para" id="N65539">The 20-question Revised Study Process Questionnaire (R-SPQ-2F), which is frequently used to categorize student learning approaches as either <i>deep</i> or <i>surface</i>, was administered to three sections of Anatomy &amp; Physiology (A&amp;P) courses at a highest research university in the southeastern United States as part of a larger research project. Two hundred thirty-one (231) respondents completed the full survey and 11 participants were recruited to a comparative case study. Initial review of interview transcripts raised concerns about the validity of the R-SPQ-2F results with the population of interest. Interview transcripts were coded using <i>a priori</i> codes corresponding to the R-SPQ-2F items, and qualitative and quantitative results were then triangulated. Additional survey responses were collected in a subsequent semester and a confirmatory factor analysis (CFA) was performed using the complete responses from 381 students. The CFA yielded similar or better measures of reliability and fit to the two-factor structure as those in previously reported work by other authors. Nonetheless, findings from triangulation suggest that the R-SPQ-2F was not able to group students by deep and surface approaches to learning in the context of an undergraduate A&amp;P course. In addition, six interviews (3 deep, 3 surface) demonstrated a new theme of <i>surface leading to deep</i> with participants indicating that memorization was necessary for the purpose of gaining a full understanding of the course material. This mixed method analysis calls into question whether the results are valid for separating student approaches into the previously published descriptions of deep and surface approaches. The finding of the <i>surface leading to deep</i> orientation, which may align with previous descriptions of an <i>achieving</i> approach, has significant implications for both research and instruction, as memorizing and other “surface” strategies are often minimized and discouraged, yet are an important step in student learning.</p>]]></description>
            <pubDate><![CDATA[2021-04-29T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Using survival prediction techniques to learn consumer-specific reservation price distributions]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766068319940-96f3fd33-95ac-41ac-806e-73a414ffea32/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249182</link>
            <description><![CDATA[<p class="para" id="N65539">A consumer’s “reservation price” (RP) is the highest price that s/he is willing to pay for one unit of a specified product or service. It is an essential concept in many applications, including personalized pricing, auction and negotiation. While consumers will not volunteer their RPs, we may be able to predict these values, based on each consumer’s specific information, using a model learned from earlier consumer transactions. Here, we view each such (non)transaction as a <i>censored observation</i>, which motivates us to use techniques from survival analysis/prediction, to produce models that can generate a consumer-specific RP distribution, based on features of each new consumer. To validate this framework of RP, we run experiments on realistic data, with four survival prediction methods. These models performed very well (under three different criteria) on the task of estimating consumer-specific RP distributions, which shows that our RP framework can be effective.</p>]]></description>
            <pubDate><![CDATA[2021-04-29T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Graph diffusion distance: Properties and efficient computation]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766067778636-3e6b96c3-82a5-4436-a3a0-6ba4e4ac9845/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249624</link>
            <description><![CDATA[<p class="para" id="N65539">We define a new family of similarity and distance measures on graphs, and explore their theoretical properties in comparison to conventional distance metrics. These measures are defined by the solution(s) to an optimization problem which attempts find a map minimizing the discrepancy between two graph Laplacian exponential matrices, under norm-preserving and sparsity constraints. Variants of the distance metric are introduced to consider such optimized maps under sparsity constraints as well as fixed time-scaling between the two Laplacians. The objective function of this optimization is multimodal and has discontinuous slope, and is hence difficult for univariate optimizers to solve. We demonstrate a novel procedure for efficiently calculating these optima for two of our distance measure variants. We present numerical experiments demonstrating that (a) upper bounds of our distance metrics can be used to distinguish between lineages of related graphs; (b) our procedure is faster at finding the required optima, by as much as a factor of 10<sup>3</sup>; and (c) the upper bounds satisfy the triangle inequality exactly under some assumptions and approximately under others. We also derive an upper bound for the distance between two graph products, in terms of the distance between the two pairs of factors. Additionally, we present several possible applications, including the construction of infinite “graph limits” by means of Cauchy sequences of graphs related to one another by our distance measure.</p>]]></description>
            <pubDate><![CDATA[2021-04-27T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Accurate contact-based modelling of repeat proteins predicts the structure of new repeats protein families]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766067094069-18d283ac-f203-4e4b-956a-7999546b0df2/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pcbi.1008798</link>
            <description><![CDATA[<p class="para" id="N65539">Repeat proteins are abundant in eukaryotic proteomes. They are involved in many eukaryotic specific functions, including signalling. For many of these proteins, the structure is not known, as they are difficult to crystallise. Today, using direct coupling analysis and deep learning it is often possible to predict a protein’s structure. However, the unique sequence features present in repeat proteins have been a challenge to use direct coupling analysis for predicting contacts. Here, we show that deep learning-based methods (trRosetta, DeepMetaPsicov (DMP) and PconsC4) overcomes this problem and can predict intra- and inter-unit contacts in repeat proteins. In a benchmark dataset of 815 repeat proteins, about 90% can be correctly modelled. Further, among 48 PFAM families lacking a protein structure, we produce models of forty-one families with estimated high accuracy.</p><p class="para" id="N65542">Repeat proteins are widespread among organisms and particularly abundant in eukaryotic proteomes. Their primary sequence presents repetition in the amino acid sequences that origin structures with repeated folds/domains. Although the repeated units often can be recognised from the sequence alone, often structural information is missing. Here, we used contact prediction for predicting the structure of repeats protein directly from their primary sequences. We benchmark the methods on a dataset comprehensive of all the known repeated structures. We evaluate the contact predictions and the obtained models for different classes of repeat proteins. Further, we develop and benchmark a quality assessment (QA) method specific for repeat proteins. Finally, we used the prediction pipeline for all PFAM repeat families without resolved structures and found that forty-one of them could be modelled with high accuracy.</p>]]></description>
            <pubDate><![CDATA[2021-04-15T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Increasing prediction accuracy of pathogenic staging by sample augmentation with a GAN]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766066478779-c8fa935e-23bc-4a02-9137-f8f128e9e376/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0250458</link>
            <description><![CDATA[<p class="para" id="N65539">Accurate prediction of cancer stage is important in that it enables more appropriate treatment for patients with cancer. Many measures or methods have been proposed for more accurate prediction of cancer stage, but recently, machine learning, especially deep learning-based methods have been receiving increasing attention, mostly owing to their good prediction accuracy in many applications. Machine learning methods can be applied to high throughput DNA mutation or RNA expression data to predict cancer stage. However, because the number of genes or markers generally exceeds 10,000, a considerable number of data samples is required to guarantee high prediction accuracy. To solve this problem of a small number of clinical samples, we used a Generative Adversarial Networks (GANs) to augment the samples. Because GANs are not effective with whole genes, we first selected significant genes using DNA mutation data and random forest feature ranking. Next, RNA expression data for selected genes were expanded using GANs. We compared the classification accuracies using original dataset and expanded datasets generated by proposed and existing methods, using random forest, Deep Neural Networks (DNNs), and 1-Dimensional Convolutional Neural Networks (1DCNN). When using the 1DCNN, the F1 score of GAN5 (a 5-fold increase in data) was improved by 39% in relation to the original data. Moreover, the results using only 30% of the data were better than those using all of the data. Our attempt is the first to use GAN for augmentation using numeric data for both DNA and RNA. The augmented datasets obtained using the proposed method demonstrated significantly increased classification accuracy for most cases. By using GAN and 1DCNN in the prediction of cancer stage, we confirmed that good results can be obtained even with small amounts of samples, and it is expected that a great deal of the cost and time required to obtain clinical samples will be reduced. The proposed sample augmentation method could also be applied for other purposes, such as prognostic prediction or cancer classification.</p>]]></description>
            <pubDate><![CDATA[2021-04-27T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[A fused-image-based approach to detect obstructive sleep apnea using a single-lead ECG and a 2D convolutional neural network]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766066392300-6720a34c-6b4b-41cf-a463-a836fa34111a/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0250618</link>
            <description><![CDATA[<p class="para" id="N65539">Obstructive sleep apnea (OSA) is a common chronic sleep disorder that disrupts breathing during sleep and is associated with many other medical conditions, including hypertension, coronary heart disease, and depression. Clinically, the standard for diagnosing OSA involves nocturnal polysomnography (PSG). However, this requires expert human intervention and considerable time, which limits the availability of OSA diagnosis in public health sectors. Therefore, electrocardiogram (ECG)-based methods for OSA detection have been proposed to automate the polysomnography procedure and reduce its discomfort. So far, most of the proposed approaches rely on feature engineering, which calls for advanced expert knowledge and experience. This paper proposes a novel fused-image-based technique that detects OSA using only a single-lead ECG signal. In the proposed approach, a convolutional neural network extracts features automatically from images created with one-minute ECG segments. The proposed network comprises 37 layers, including four residual blocks, a dense layer, a dropout layer, and a soft-max layer. In this study, three time–frequency representations, namely the scalogram, the spectrogram, and the Wigner–Ville distribution, were used to investigate the effectiveness of the fused-image-based approach. We found that blending scalogram and spectrogram images further improved the system’s discriminative characteristics. Seventy ECG recordings from the PhysioNet Apnea-ECG database were used to train and evaluate the proposed model using 10-fold cross validation. The results of this study demonstrated that the proposed classifier can perform OSA detection with an average accuracy, recall, and specificity of 92.4%, 92.3%, and 92.6%, respectively, for the fused spectral images.</p>]]></description>
            <pubDate><![CDATA[2021-04-26T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The phase space of meaning model of psychopathology: A computer simulation modelling study]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766065904346-d253179c-fbe2-43e1-bf68-855327f2746d/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249320</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Introduction</h3><p class="para" id="N65543">The hypothesis of a general psychopathology factor that underpins all common forms of mental disorders has been gaining momentum in contemporary clinical research and is known as the <i>p</i> factor hypothesis. Recently, a semiotic, embodied, and psychoanalytic conceptualisation of the <i>p</i> factor has been proposed called the Harmonium Model, which provides a computational account of such a construct. This research tested the core tenet of the Harmonium model, which is the idea that psychopathology can be conceptualised as due to poorly-modulable cognitive processes, and modelled the concept of Phase Space of Meaning (PSM) at the computational level.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Method</h3><p class="para" id="N65555">Two studies were performed, both based on a simulation design implementing a deep learning model, simulating a cognitive process: a classification task. The level of performance of the task was considered the simulated equivalent to the normality-psychopathology continuum, the dimensionality of the neural network’s internal computational dynamics being the simulated equivalent of the PSM’s dimensionality.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Results</h3><p class="para" id="N65561">The neural networks’ level of performance was shown to be associated with the characteristics of the internal computational dynamics, assumed to be the simulated equivalent of poorly-modulable cognitive processes.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Discussion</h3><p class="para" id="N65567">Findings supported the hypothesis. They showed that the neural network’s low performance was a matter of the combination of predicted characteristics of the neural networks’ internal computational dynamics. Implications, limitations, and further research directions are discussed.</p></div>]]></description>
            <pubDate><![CDATA[2021-04-26T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The role of foreign technologies and R&amp;D in innovation processes within catching-up CEE countries]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766065036281-5f8cef6d-2814-4678-92d7-305138370813/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0250307</link>
            <description><![CDATA[<p class="para" id="N65539">Prior research showed that there is a growing consensus among researchers, which point out a key role of external knowledge sources such as external R&amp;D and technologies in enhancing firms´ innovation. However, firms´ from catching-up Central and Eastern European (CEE) countries have already shown in the past that their innovation models differ from those applied, for example, in Western Europe. This study therefore introduces a novel two-staged model combining artificial neural networks and random forests to reveal the importance of internal and external factors influencing firms´ innovation performance in the case of 3,361 firms from six catching-up CEE countries (Czech Republic, Slovakia, Poland, Estonia, Latvia and Lithuania), by using the World Banks´ Enterprise Survey data from 2019. We confirm the hypothesis that innovators in the catching-up CEE countries depend more on internal knowledge sources and, moreover, that participation in the firms groups represents an important factor of firms´ innovation. Surprisingly, we reject the hypothesis that foreign technologies are a crucial source of external knowledge. This study contributes to the theories of open innovation and absorptive capacity in the context of selected CEE countries and provides several practical implications for firms.</p>]]></description>
            <pubDate><![CDATA[2021-04-22T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Sparse Poisson regression via mixed-integer optimization]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766064825418-5aa6fda4-edc0-46ba-8edf-b2344c1ac811/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249916</link>
            <description><![CDATA[<p class="para" id="N65539">We present a mixed-integer optimization (MIO) approach to sparse Poisson regression. The MIO approach to sparse linear regression was first proposed in the 1970s, but has recently received renewed attention due to advances in optimization algorithms and computer hardware. In contrast to many sparse estimation algorithms, the MIO approach has the advantage of finding the best subset of explanatory variables with respect to various criterion functions. In this paper, we focus on a sparse Poisson regression that maximizes the weighted sum of the log-likelihood function and the <i>L</i><sub>2</sub>-regularization term. For this problem, we derive a mixed-integer quadratic optimization (MIQO) formulation by applying a piecewise-linear approximation to the log-likelihood function. Optimization software can solve this MIQO problem to optimality. Moreover, we propose two methods for selecting a limited number of tangent lines effective for piecewise-linear approximations. We assess the efficacy of our method through computational experiments using synthetic and real-world datasets. Our methods provide better log-likelihood values than do conventional greedy algorithms in selecting tangent lines. In addition, our MIQO formulation delivers better out-of-sample prediction performance than do forward stepwise selection and <i>L</i><sub>1</sub>-regularized estimation, especially in low-noise situations.</p>]]></description>
            <pubDate><![CDATA[2021-04-22T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Artificial intelligence based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsa<sup>a</sup>)]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766064488380-9f1530db-7eb5-49f3-bddb-11ef26bad440/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249769</link>
            <description><![CDATA[<p class="para" id="N65539">The Dead Sea Scrolls are tangible evidence of the Bible’s ancient scribal culture. This study takes an innovative approach to palaeography—the study of ancient handwriting—as a new entry point to access this scribal culture. One of the problems of palaeography is to determine writer identity or difference when the writing style is near uniform. This is exemplified by the Great Isaiah Scroll (1QIsa<sup>a</sup>). To this end, we use pattern recognition and artificial intelligence techniques to innovate the palaeography of the scrolls and to pioneer the microlevel of individual scribes to open access to the Bible’s ancient scribal culture. We report new evidence for a breaking point in the series of columns in this scroll. Without prior assumption of writer identity, based on point clouds of the reduced-dimensionality feature-space, we found that columns from the first and second halves of the manuscript ended up in two distinct zones of such scatter plots, notably for a range of digital palaeography tools, each addressing very different featural aspects of the script samples. In a secondary, independent, analysis, now assuming writer difference and using yet another independent feature method and several different types of statistical testing, a switching point was found in the column series. A clear phase transition is apparent in columns 27–29. We also demonstrated a difference in distance variances such that the variance is higher in the second part of the manuscript. Given the statistically significant differences between the two halves, a tertiary, post-hoc analysis was performed using visual inspection of character heatmaps and of the most discriminative Fraglet sets in the script. Demonstrating that two main scribes, each showing different writing patterns, were responsible for the Great Isaiah Scroll, this study sheds new light on the Bible’s ancient scribal culture by providing new, tangible evidence that ancient biblical texts were not copied by a single scribe only but that multiple scribes, while carefully mirroring another scribe’s writing style, could closely collaborate on one particular manuscript.</p>]]></description>
            <pubDate><![CDATA[2021-04-21T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The influence of algorithms on political and dating decisions]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766063501510-2530e5b8-4bdc-4713-aa05-5e8e297e13ac/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249454</link>
            <description><![CDATA[<p class="para" id="N65539">Artificial intelligence algorithms are ubiquitous in daily life, and this is motivating the development of some institutional initiatives to ensure trustworthiness in Artificial Intelligence (AI). However, there is not enough research on how these algorithms can influence people’s decisions and attitudes. The present research examines whether algorithms can persuade people, explicitly or covertly, on whom to vote and date, or whether, by contrast, people would reject their influence in an attempt to confirm their personal freedom and independence. In four experiments, we found that persuasion was possible and that different styles of persuasion (e.g., explicit, covert) were more effective depending on the decision context (e.g., political and dating). We conclude that it is important to educate people against trusting and following the advice of algorithms blindly. A discussion on who owns and can use the data that makes these algorithms work efficiently is also necessary.</p>]]></description>
            <pubDate><![CDATA[2021-04-21T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Perception and prediction of the putting distance of robot putting movements under different visual/viewing conditions]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766062940409-aa3a45ca-beea-49f1-a25d-69c030be169a/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249518</link>
            <description><![CDATA[<p class="para" id="N65539">The purpose of this paper is to examine, whether and under which conditions humans are able to predict the putting distance of a robotic device. Based on the “flash-lag effect” (FLE) it was expected that the prediction errors increase with increasing putting velocity. Furthermore, we hypothesized that the predictions are more accurate and more confident if human observers operate under full vision (F-RCHB) compared to either temporal occlusion (I-RCHB) or spatial occlusion (invisible ball, F-RHC, or club, F-B). In two experiments, 48 video sequences of putt movements performed by a BioRob robot arm were presented to thirty-nine students (age: 24.49±3.20 years). In the experiments, video sequences included six putting distances (1.5, 2.0, 2.5, 3.0, 3.5, and 4.0 m; experiment 1) under full versus incomplete vision (F-RCHB versus I-RCHB) and three putting distances (2. 0, 3.0, and 4.0 m; experiment 2) under the four visual conditions (F-RCHB, I-RCHB, F-RCH, and F-B). After the presentation of each video sequence, the participants estimated the putting distance on a scale from 0 to 6 m and provided their confidence of prediction on a 5-point scale. Both experiments show comparable results for the respective dependent variables (error and confidence measures). The participants consistently overestimated the putting distance under the full vision conditions; however, the experiments did not show a pattern that was consistent with the FLE. Under the temporal occlusion condition, a prediction was not possible; rather a random estimation pattern was found around the centre of the prediction scale (3 m). Spatial occlusion did not affect errors and confidence of prediction. The experiments indicate that temporal constraints seem to be more critical than spatial constraints. The FLE may not apply to distance prediction compared to location estimation.</p>]]></description>
            <pubDate><![CDATA[2021-04-23T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[A novel approach to dry weight adjustments for dialysis patients using machine learning]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766062747660-0cb082f5-c916-478b-9167-c0c51fe5fb28/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0250467</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Background and aims</h3><p class="para" id="N65543">Knowledge of the proper dry weight plays a critical role in the efficiency of dialysis and the survival of hemodialysis patients. Recently, bioimpedance spectroscopy(BIS) has been widely used for set dry weight in hemodialysis patients. However, BIS is often misrepresented in clinical healthy weight. In this study, we tried to predict the clinically proper dry weight (DW<sub>CP</sub>) using machine learning for patient’s clinical information including BIS. We then analyze the factors that influence the prediction of the clinical dry weight.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Methods</h3><p class="para" id="N65552">As a retrospective, single center study, data of 1672 hemodialysis patients were reviewed. DW<sub>CP</sub> data were collected when the dry weight was measured using the BIS (DW<sub>BIS</sub>). The gap between the two (Gap<sub>DW</sub>) was calculated and then grouped and analyzed based on gaps of 1 kg and 2 kg.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Results</h3><p class="para" id="N65567">Based on the gap between DW<sub>BIS</sub> and DW<sub>CP</sub>, 972, 303, and 384 patients were placed in groups with gaps of &lt;1 kg, ≧1kg and &lt;2 kg, and ≧2 kg, respectively. For less than 1 kg and 2 kg of GapDW, It can be seen that the average accuracies for the two groups are 83% and 72%, respectively, in usign XGBoost machine learning. As Gap<sub>DW</sub> increases, it is more difficult to predict the target property. As Gap<sub>DW</sub> increase, the mean values of hemoglobin, total protein, serum albumin, creatinine, phosphorus, potassium, and the fat tissue index tended to decrease. However, the height, total body water, extracellular water (ECW), and ECW to intracellular water ratio tended to increase.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Conclusions</h3><p class="para" id="N65585">Machine learning made it slightly easier to predict DW<sub>CP</sub> based on DW<sub>BIS</sub> under limited conditions and gave better insights into predicting DW<sub>CP</sub>. Malnutrition-related factors and ECW were important in reflecting the differences between DW<sub>BIS</sub> and DW<sub>CP</sub>.</p></div>]]></description>
            <pubDate><![CDATA[2021-04-23T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[eNose-TB: A trial study protocol of electronic nose for tuberculosis screening in Indonesia]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766062125724-c7fa476e-72eb-46b8-927b-8c68716ffaab/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249689</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Background</h3><p class="para" id="N65543">Even though conceptually, Tuberculosis (TB) is almost always curable, it is currently the world’s leading infectious killer. Patients with pulmonary TB are the source of transmission. Approximately 23% of the world’s population is believed to be latently infected with TB bacteria, and 5–15% of them will progress at any point in time to develop the disease. There was a global diagnostic gap of 2.9 million between notifications of new cases and the estimated number of incident cases, and Indonesia carries the third-highest of this gap. Therefore, screening TB among the community is of great importance to prevent further transmission and infection. The electronic nose for screening TB (eNose-TB) project is initiated in Yogyakarta, Indonesia, to screen TB by breath test with an electronic-nose that is easy-to-use, point-of-care, does not expose patients to radiation, and can be produced at low cost.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Methods/Design</h3><p class="para" id="N65549">The objectives of the two-phase planned project are to: 1) investigate the potential of an eNose-TB as a screening tool in Indonesia, in comparison with screening with clinical symptoms and chest radiology, which are currently used as a standard, and 2) analyze the time and cost of a screening algorithm with eNose-TB to obtain additional case detection. A cross-sectional study will be conducted in the first phase to validate the eNose-TB. The validation phase will involve 395 presumptive TB patients in the Surakarta General Hospital, Central Java. In the second phase, a cross-sectional research will be conducted, involving 1,383 adults and children in the municipality of Yogyakarta and Kulon Progo district of Yogyakarta Province.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Discussion</h3><p class="para" id="N65555">The findings will provide data concerning the sensitivity and specificity of the eNose-TB as a screening tool for tuberculosis, and the time and cost analysis of a screening algorithm with the eNose.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Trial registration</h3><p class="para" id="N65561">NCT04567498; https://clinicaltrials.gov/.</p></div>]]></description>
            <pubDate><![CDATA[2021-04-21T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[An intelligent cluster optimization algorithm based on Whale
Optimization Algorithm for VANETs (WOACNET)]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766062041620-c45d58e4-1d2e-4b44-9334-66d9caca109f/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0250271</link>
            <description><![CDATA[<p class="para" id="N65539">Vehicular Ad hoc Networks (VANETs) an important category in networking focuses on
many applications, such as safety and intelligent traffic management systems.
The high node mobility and sparse vehicle distribution (on the road) compromise
VANETs network scalability and rapid topology, hence creating major challenges,
such as network physical layout formation, unstable links to enable robust,
reliable, and scalable vehicle communication, especially in a dense traffic
network. This study discusses a novel optimization approach considering
transmission range, node density, speed, direction, and grid size during
clustering. Whale Optimization Algorithm for Clustering in Vehicular Ad hoc
Networks (WOACNET) was introduced to select an optimum cluster head (CH) and was
calculated and evaluated based on intelligence and capability. Initially,
simulations were performed, Subsequently, rigorous experimentations were
conducted on WOACNET. The model was compared and evaluated with state-of-the-art
well-established other methods, such as Gray Wolf Optimization (GWO) and Ant
Lion Optimization (ALO) employing various performance metrics. The results
demonstrate that the developed method performance is well ahead compared to
other methods in VANET in terms of cluster head, varying transmission ranges,
grid size, and nodes. The developed method results in achieving an overall 46%
enhancement in cluster optimization and an F-value of 31.64 compared to other
established methods (11.95 and 22.50) consequently, increase in cluster
lifetime.</p>]]></description>
            <pubDate><![CDATA[2021-04-21T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Crowdsourcing airway annotations in chest computed tomography images]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766061862668-17b710a1-ba68-4347-b4a9-3ec980717ec7/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249580</link>
            <description><![CDATA[<p class="para" id="N65539">Measuring airways in chest computed tomography (CT) scans is important for characterizing diseases such as cystic fibrosis, yet very time-consuming to perform manually. Machine learning algorithms offer an alternative, but need large sets of annotated scans for good performance. We investigate whether crowdsourcing can be used to gather airway annotations. We generate image slices at known locations of airways in 24 subjects and request the crowd workers to outline the airway lumen and airway wall. After combining multiple crowd workers, we compare the measurements to those made by the experts in the original scans. Similar to our preliminary study, a large portion of the annotations were excluded, possibly due to workers misunderstanding the instructions. After excluding such annotations, moderate to strong correlations with the expert can be observed, although these correlations are slightly lower than inter-expert correlations. Furthermore, the results across subjects in this study are quite variable. Although the crowd has potential in annotating airways, further development is needed for it to be robust enough for gathering annotations in practice. For reproducibility, data and code are available online: http://github.com/adriapr/crowdairway.git.</p>]]></description>
            <pubDate><![CDATA[2021-04-22T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[<span style="font-variant: all-small-caps">TriatoDex</span>, an electronic identification key to the Triatominae (Hemiptera: Reduviidae), vectors of Chagas disease: Development, description, and performance]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766061850240-2af1e082-017f-4b2c-a408-87ce6f9ff974/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0248628</link>
            <description><![CDATA[<p class="para" id="N65539">Correct identification of triatomine bugs is crucial for Chagas disease surveillance, yet available taxonomic keys are outdated, incomplete, or both. Here we present <span style="font-variant: all-small-caps">TriatoDex</span>, an Android app-based pictorial, annotated, polytomous key to the Triatominae. <span style="font-variant: all-small-caps">TriatoDex</span> was developed using Android Studio and tested by 27 Brazilian users. Each user received a box with pinned, number-labeled, adult triatomines (33 species in total) and was asked to identify each bug to the species level. We used generalized linear mixed models (with user- and species-ID random effects) and information-theoretic model evaluation/averaging to investigate <span style="font-variant: all-small-caps">TriatoDex</span> performance. <span style="font-variant: all-small-caps">TriatoDex</span> encompasses 79 questions and 554 images of the 150 triatomine-bug species described worldwide up to 2017. <span style="font-variant: all-small-caps">TriatoDex</span>-based identification was correct in 78.9% of 824 tasks. <span style="font-variant: all-small-caps">TriatoDex</span> performed better in the hands of trained taxonomists (93.3% <i>vs</i>. 72.7% correct identifications; model-averaged, adjusted odds ratio 5.96, 95% confidence interval [CI] 3.09–11.48). In contrast, user age, gender, primary job (including academic research/teaching or disease surveillance), workplace (including universities, a reference laboratory for triatomine-bug taxonomy, or disease-surveillance units), and basic training (from high school to biology) all had negligible effects on <span style="font-variant: all-small-caps">TriatoDex</span> performance. Our analyses also suggest that, as <span style="font-variant: all-small-caps">TriatoDex</span> results accrue to cover more taxa, they may help pinpoint triatomine-bug species that are consistently harder (than average) to identify. In a pilot comparison with a standard, printed key (370 tasks by seven users), <span style="font-variant: all-small-caps">TriatoDex</span> performed similarly (84.5% correct assignments, CI 68.9–94.0%), but identification was 32.8% (CI 24.7–40.1%) faster on average–for a mean absolute saving of ~2.3 minutes per bug-identification task. <span style="font-variant: all-small-caps">TriatoDex</span> holds much promise as a handy, flexible, and reliable tool for triatomine-bug identification; an updated iOS/Android version is under development. We expect that, with continuous refinement derived from evolving knowledge and user feedback, <span style="font-variant: all-small-caps">TriatoDex</span> will substantially help strengthen both entomological surveillance and research on Chagas disease vectors.</p>]]></description>
            <pubDate><![CDATA[2021-04-22T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Clustered embedding using deep learning to analyze urban mobility based on complex transportation data]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766060337530-1d840815-a724-41cc-ad9c-f646ed96078b/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249318</link>
            <description><![CDATA[<p class="para" id="N65539">Urban mobility is a vital aspect of any city and often influences its physical shape as well as its level of economic and social development. A thorough analysis of mobility patterns in urban areas can provide various benefits, such as the prediction of traffic flow and public transportation usage. In particular, based on its exceptional ability to extract patterns from complex large-scale data, embedding based on deep learning is a promising method for analyzing the mobility patterns of urban residents. However, as urban mobility becomes increasingly complex, it becomes difficult to embed patterns into a single vector because of its limited capacity. In this paper, we propose a novel method for analyzing urban mobility based on deep learning. The proposed method involves clustering mobility patterns and embedding them to capture their implicit meaning. Clustering groups mobility patterns based on their spatiotemporal characteristics, and embedding provides meaningful information regarding both individual residents (i.e., personalized mobility) and all residents as a whole, enabling a more effective analysis of mobility patterns. Experiments were performed to predict the successive points of interest (POIs) based on transportation data collected from 1.5 million citizens in a large metropolitan city; the results demonstrate that the proposed method achieves top-1, 3, and 5 accuracies of 73.64%, 88.65%, and 91.54%, respectively, which are much higher than those of the conventional method (59.48%, 75.85%, and 80.1%, respectively). We also demonstrate that the proposed method facilitates the analysis of urban mobility through arithmetic operations between POI vectors.</p>]]></description>
            <pubDate><![CDATA[2021-04-20T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Combining natural language processing and metabarcoding to reveal pathogen-environment associations]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766053454257-a23f4677-3cc5-4420-b792-00b88078d53c/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pntd.0008755</link>
            <description><![CDATA[<p class="para" id="N65539"><i>Cryptococcus neoformans</i> is responsible for life-threatening infections that primarily affect immunocompromised individuals and has an estimated worldwide burden of 220,000 new cases each year—with 180,000 resulting deaths—mostly in sub-Saharan Africa. Surprisingly, little is known about the ecological niches occupied by <i>C</i>. <i>neoformans</i> in nature. To expand our understanding of the distribution and ecological associations of this pathogen we implement a Natural Language Processing approach to better describe the niche of <i>C</i>. <i>neoformans</i>. We use a Latent Dirichlet Allocation model to <i>de novo</i> topic model sets of metagenetic research articles written about varied subjects which either explicitly mention, inadvertently find, or fail to find <i>C</i>. <i>neoformans</i>. These articles are all linked to NCBI Sequence Read Archive datasets of 18S ribosomal RNA and/or Internal Transcribed Spacer gene-regions. The number of topics was determined based on the model coherence score, and articles were assigned to the created topics via a Machine Learning approach with a Random Forest algorithm. Our analysis provides support for a previously suggested linkage between <i>C</i>. <i>neoformans</i> and soils associated with decomposing wood. Our approach, using a search of single-locus metagenetic data, gathering papers connected to the datasets, <i>de novo</i> determination of topics, the number of topics, and assignment of articles to the topics, illustrates how such an analysis pipeline can harness large-scale datasets that are published/available but not necessarily fully analyzed, or whose metadata is not harmonized with other studies. Our approach can be applied to a variety of systems to assert potential evidence of environmental associations.</p><p class="para" id="N65542">We expand the utility of Natural Language Processing (NLP), backtracking through metabarcodes, utilizing papers that may not mention our subject of interest, <i>C</i>. <i>neoformans</i>, in a departure from usual text analysis methods. We confirm that <i>C</i>. <i>neoformans</i> is associated with decomposing wood which is reinforced by the inferred literature studied here on <i>C</i>. <i>neoformans</i> and its close congeneric relatives. This work demonstrates the potential utility of pairing NLP with single-locus metagenetic data for the study of Neglected Tropical Diseases. While the results of this article are largely confirmatory, we present a novel method to study the ecological niches of rare pathogens that leverages the immense amount of data available to researchers in the NCBI Sequence Read Archive (SRA) combined with a text-mining analysis based on Natural Language Processing. We demonstrate that text processing, noun identification, and verb identification can play an important role in analyzing a large corpus of documents together with metagenetic data. Forging this connection requires access to all of the available ecological 18S ribosomal RNA and Internal Transcribed Spacer NCBI SRA datasets. These datasets use metabarcoding to query taxonomic diversity in eukaryotic organisms, and in the case of the Internal Transcribed Spacer, they specifically target Fungi. The presence of specific species is inferred when diagnostic 18S or ITS gene region sequences are found in the SRA data. We searched for <i>C</i>. <i>neoformans</i> in all 18S and ITS datasets available and gathered all associated journal articles that either cite the SRA data accessions or are cited in the SRA data accessions. Published metagenetic data often have associated metadata including: latitude and longitude, temperature, and other physical characteristics describing the conditions in which the metagenetic sample was collected. These metadata are not always presented in consistent formats, so harmonizing study methods may be needed to appropriately compare metagenetic data as commonly required in metanalysis studies. We present an analysis which takes as input articles associated with SRA datasets that were found to contain evidence of <i>C</i>. <i>neoformans</i>. We apply NLP methods to this corpus of articles to describe the niche of <i>C</i>. <i>neoformans</i>. Our results reinforce the current understanding of <i>C</i>. <i>neoformans</i>’s niche, indicating the pertinence of employing an NLP analysis to identify the niche of an organism. This approach could further the description of virtually any other organism that routinely appears in metagenetic surveys, especially pathogens, whose ecological niches are unknown or poorly understood.</p>]]></description>
            <pubDate><![CDATA[2021-04-07T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Emulative, coherent, and causal dynamics between large-scale brain networks are neurobiomarkers of Accelerated Cognitive Ageing in epilepsy]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766047081161-45666cf0-6395-4c38-bcd6-f7ff32fac459/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0250222</link>
            <description><![CDATA[<p class="para" id="N65539">Accelerated cognitive ageing (ACA) is an ageing co-morbidity in epilepsy that is diagnosed through the observation of an evident IQ decline of more than 1 standard deviation (15 points) around the age of 50 years old. To understand the mechanism of action of this pathology, we assessed brain dynamics with the use of resting-state fMRI data. In this paper, we present novel and promising methods to extract brain dynamics between large-scale resting-state networks: the emulative power, wavelet coherence, and granger causality between the networks were extracted in two resting-state sessions of 24 participants (10 ACA, 14 controls). We also calculated the widely used static functional connectivity to compare the methods. To find the best biomarkers of ACA, and have a better understanding of this epilepsy co-morbidity we compared the aforementioned between-network neurodynamics using classifiers and known machine learning algorithms; and assessed their performance. Results show that features based on the evolutionary game theory on networks approach, the emulative powers, are the best descriptors of the co-morbidity, using dynamics associated with the default mode and dorsal attention networks. With these dynamic markers, linear discriminant analysis could identify ACA patients at 82.9% accuracy. Using wavelet coherence features with decision-tree algorithm, and static functional connectivity features with support vector machine, ACA could be identified at 77.1% and 77.9% accuracy respectively. Granger causality fell short of being a relevant biomarker with best classifiers having an average accuracy of 67.9%. Combining the features based on the game theory, wavelet coherence, Granger-causality, and static functional connectivity- approaches increased the classification performance up to 90.0% average accuracy using support vector machine with a peak accuracy of 95.8%. The dynamics of the networks that lead to the best classifier performances are known to be challenged in elderly. Since our groups were age-matched, the results are in line with the idea of ACA patients having an accelerated cognitive decline. This classification pipeline is promising and could help to diagnose other neuropsychiatric disorders, and contribute to the field of psychoradiology.</p>]]></description>
            <pubDate><![CDATA[2021-04-16T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Behavioral discrimination and time-series phenotyping of birdsong performance]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766046959978-d3db3c3c-f584-4f18-92a4-8f1c1fd41c94/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pcbi.1008820</link>
            <description><![CDATA[<p class="para" id="N65539">Variation in the acoustic structure of vocal signals is important to communicate social information. However, relatively little is known about the features that receivers extract to decipher relevant social information. Here, we took an expansive, bottom-up approach to delineate the feature space that could be important for processing social information in zebra finch song. Using operant techniques, we discovered that female zebra finches can consistently discriminate brief song phrases (“motifs”) from different social contexts. We then applied machine learning algorithms to classify motifs based on thousands of time-series features and to uncover acoustic features for motif discrimination. In addition to highlighting classic acoustic features, the resulting algorithm revealed novel features for song discrimination, for example, measures of time irreversibility (i.e., the degree to which the statistical properties of the actual and time-reversed signal differ). Moreover, the algorithm accurately predicted female performance on individual motif exemplars. These data underscore and expand the promise of broad time-series phenotyping to acoustic analyses and social decision-making.</p><p class="para" id="N65542">Variation in vocal performance can provide information about the social context and motivation of the signaler. For example, prosodic changes to the pitch, tempo, or loudness of speech can reveal a speaker’s emotional state or motivation, even when speech content is unchanged. Similarly, in songbirds like the zebra finch, males produce songs with the same acoustic elements but with different vocal performance when courting females and when singing alone, and females strongly prefer to hear the courtship song. Here we integrated behavioral and computational approaches to reveal the acoustic features that female zebra finches might use for social discrimination. We first discovered that females excelled at distinguishing between brief phrases of courtship and non-courtship song. We next extracted thousands of time-series features from courtship and non-courtship phrases using a highly-comparative time series analysis (HCTSA) toolbox, and trained machine learning algorithms to use those features to discriminate between phrases. The machine learning algorithm identified features important for discriminating between courtship and non-courtship song phrases, some of which have not been implicated before, and reliably predicted the discrimination abilities of females. Together, these data highlight the power of expansive and bottom-up approaches to reveal acoustic features important for social discrimination.</p>]]></description>
            <pubDate><![CDATA[2021-04-08T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Predicting breast cancer 5-year survival using machine learning: A systematic review]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766046210637-4d6297ee-cbea-463c-99a6-4fb850f6748a/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0250370</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Background</h3><p class="para" id="N65543">Accurately predicting the survival rate of breast cancer patients is a major issue for cancer researchers. Machine learning (ML) has attracted much attention with the hope that it could provide accurate results, but its modeling methods and prediction performance remain controversial. The aim of this systematic review is to identify and critically appraise current studies regarding the application of ML in predicting the 5-year survival rate of breast cancer.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Methods</h3><p class="para" id="N65549">In accordance with the PRISMA guidelines, two researchers independently searched the PubMed (including MEDLINE), Embase, and Web of Science Core databases from inception to November 30, 2020. The search terms included breast neoplasms, survival, machine learning, and specific algorithm names. The included studies related to the use of ML to build a breast cancer survival prediction model and model performance that can be measured with the value of said verification results. The excluded studies in which the modeling process were not explained clearly and had incomplete information. The extracted information included literature information, database information, data preparation and modeling process information, model construction and performance evaluation information, and candidate predictor information.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Results</h3><p class="para" id="N65555">Thirty-one studies that met the inclusion criteria were included, most of which were published after 2013. The most frequently used ML methods were decision trees (19 studies, 61.3%), artificial neural networks (18 studies, 58.1%), support vector machines (16 studies, 51.6%), and ensemble learning (10 studies, 32.3%). The median sample size was 37256 (range 200 to 659820) patients, and the median predictor was 16 (range 3 to 625). The accuracy of 29 studies ranged from 0.510 to 0.971. The sensitivity of 25 studies ranged from 0.037 to 1. The specificity of 24 studies ranged from 0.008 to 0.993. The AUC of 20 studies ranged from 0.500 to 0.972. The precision of 6 studies ranged from 0.549 to 1. All of the models were internally validated, and only one was externally validated.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Conclusions</h3><p class="para" id="N65561">Overall, compared with traditional statistical methods, the performance of ML models does not necessarily show any improvement, and this area of research still faces limitations related to a lack of data preprocessing steps, the excessive differences of sample feature selection, and issues related to validation. Further optimization of the performance of the proposed model is also needed in the future, which requires more standardization and subsequent validation.</p></div>]]></description>
            <pubDate><![CDATA[2021-04-16T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Diagnostic performance of artificial intelligence model for pneumonia from chest radiography]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766046164165-526a628f-1512-4abd-800d-e224ba7c4600/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249399</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Objective</h3><p class="para" id="N65543">The chest X-ray (CXR) is the most readily available and common imaging modality for the assessment of pneumonia. However, detecting pneumonia from chest radiography is a challenging task, even for experienced radiologists. An artificial intelligence (AI) model might help to diagnose pneumonia from CXR more quickly and accurately. We aim to develop an AI model for pneumonia from CXR images and to evaluate diagnostic performance with external dataset.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Methods</h3><p class="para" id="N65549">To train the pneumonia model, a total of 157,016 CXR images from the National Institutes of Health (NIH) and the Korean National Tuberculosis Association (KNTA) were used (normal vs. pneumonia = 120,722 vs.36,294). An ensemble model of two neural networks with DenseNet classifies each CXR image into pneumonia or not. To test the accuracy of the models, a separate external dataset of pneumonia CXR images (n = 212) from a tertiary university hospital (Gachon University Gil Medical Center GUGMC, Incheon, South Korea) was used; the diagnosis of pneumonia was based on both the chest CT findings and clinical information, and the performance evaluated using the area under the receiver operating characteristic curve (AUC). Moreover, we tested the change of the AI probability score for pneumonia using the follow-up CXR images (7 days after the diagnosis of pneumonia, n = 100).</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Results</h3><p class="para" id="N65555">When the probability scores of the models that have a threshold of 0.5 for pneumonia, two models (models 1 and 4) having different pre-processing parameters on the histogram equalization distribution showed best AUC performances of 0.973 and 0.960, respectively. As expected, the ensemble model of these two models performed better than each of the classification models with 0.983 AUC. Furthermore, the AI probability score change for pneumonia showed a significant difference between improved cases and aggravated cases (Δ = -0.06 ± 0.14 vs. 0.06 ± 0.09, for 85 improved cases and 15 aggravated cases, respectively, <i>P</i> = 0.001) for CXR taken as a 7-day follow-up.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Conclusions</h3><p class="para" id="N65564">The ensemble model combined two different classification models for pneumonia that performed at 0.983 AUC for an external test dataset from a completely different data source. Furthermore, AI probability scores showed significant changes between cases of different clinical prognosis, which suggest the possibility of increased efficiency and performance of the CXR reading at the diagnosis and follow-up evaluation for pneumonia.</p></div>]]></description>
            <pubDate><![CDATA[2021-04-15T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Automatic image annotation for fluorescent cell nuclei segmentation]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766045761101-81d1d881-d5fb-4650-8f11-1f8b9c3408c8/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0250093</link>
            <description><![CDATA[<p class="para" id="N65539">Dataset annotation is a time and labor-intensive task and an integral requirement for training and testing deep learning models. The segmentation of images in life science microscopy requires annotated image datasets for object detection tasks such as instance segmentation. Although the amount of annotated image data has been steadily reduced due to methods such as data augmentation, the process of manual or semi-automated data annotation is the most labor and cost intensive task in the process of cell nuclei segmentation with deep neural networks. In this work we propose a system to fully automate the annotation process of a custom fluorescent cell nuclei image dataset. By that we are able to reduce nuclei labelling time by up to 99.5%. The output of our system provides high quality training data for machine learning applications to identify the position of cell nuclei in microscopy images. Our experiments have shown that the automatically annotated dataset provides coequal segmentation performance compared to manual data annotation. In addition, we show that our system enables a single workflow from raw data input to desired nuclei segmentation and tracking results without relying on pre-trained models or third-party training datasets for neural networks.</p>]]></description>
            <pubDate><![CDATA[2021-04-16T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Predicting eyes at risk for rapid glaucoma progression based on an initial visual field test using machine learning]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766045589702-ad2f3505-63fe-48de-b7ad-f9b439f59755/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249856</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Objective</h3><p class="para" id="N65543">To assess whether machine learning algorithms (MLA) can predict eyes that will undergo rapid glaucoma progression based on an initial visual field (VF) test.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Design</h3><p class="para" id="N65549">Retrospective analysis of longitudinal data.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Subjects</h3><p class="para" id="N65555">175,786 VFs (22,925 initial VFs) from 14,217 patients who completed ≥5 reliable VFs at academic glaucoma centers were included.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Methods</h3><p class="para" id="N65561">Summary measures and reliability metrics from the initial VF and age were used to train MLA designed to predict the likelihood of rapid progression. Additionally, the neural network model was trained with point-wise threshold data in addition to summary measures, reliability metrics and age. 80% of eyes were used for a training set and 20% were used as a test set. MLA test set performance was assessed using the area under the receiver operating curve (AUC). Performance of models trained on initial VF data alone was compared to performance of models trained on data from the first two VFs.</p></div><div class="section" id="sec005"><h3 class="BHead" id="nov000-5">Main outcome measures</h3><p class="para" id="N65567">Accuracy in predicting future rapid progression defined as MD worsening more than 1 dB/year.</p></div><div class="section" id="sec006"><h3 class="BHead" id="nov000-6">Results</h3><p class="para" id="N65573">1,968 eyes (8.6%) underwent rapid progression. The support vector machine model (AUC 0.72 [95% CI 0.70–0.75]) most accurately predicted rapid progression when trained on initial VF data. Artificial neural network, random forest, logistic regression and naïve Bayes classifiers produced AUC of 0.72, 0.70, 0.69, 0.68 respectively. Models trained on data from the first two VFs performed no better than top models trained on the initial VF alone. Based on the odds ratio (OR) from logistic regression and variable importance plots from the random forest model, older age (OR: 1.41 per 10 year increment [95% CI: 1.34 to 1.08]) and higher pattern standard deviation (OR: 1.31 per 5-dB increment [95% CI: 1.18 to 1.46]) were the variables in the initial VF most strongly associated with rapid progression.</p></div><div class="section" id="sec007"><h3 class="BHead" id="nov000-7">Conclusions</h3><p class="para" id="N65579">MLA can be used to predict eyes at risk for rapid progression with modest accuracy based on an initial VF test. Incorporating additional clinical data to the current model may offer opportunities to predict patients most likely to rapidly progress with even greater accuracy.</p></div>]]></description>
            <pubDate><![CDATA[2021-04-16T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Identifying geographically differentiated features of Ethopian Nile tilapia (Oreochromis niloticus) morphology with machine learning]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766036456222-ad82a194-77bd-4bc7-bd73-0297337a6959/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249593</link>
            <description><![CDATA[<p class="para" id="N65539">Visual characteristics are among the most important features for characterizing the phenotype of biological organisms. Color and geometric properties define population phenotype and allow assessing diversity and adaptation to environmental conditions. To analyze geometric properties classical morphometrics relies on biologically relevant landmarks which are manually assigned to digital images. Assigning landmarks is tedious and error prone. Predefined landmarks may in addition miss out on information which is not obvious to the human eye. The machine learning (ML) community has recently proposed new data analysis methods which by uncovering subtle features in images obtain excellent predictive accuracy. Scientific credibility demands however that results are interpretable and hence to mitigate the black-box nature of ML methods. To overcome the black-box nature of ML we apply complementary methods and investigate internal representations with saliency maps to reliably identify location specific characteristics in images of Nile tilapia populations. Analyzing fish images which were sampled from six Ethiopian lakes reveals that deep learning improves on a conventional morphometric analysis in predictive performance. A critical assessment of established saliency maps with a novel significance test reveals however that the improvement is aided by artifacts which have no biological interpretation. More interpretable results are obtained by a Bayesian approach which allows us to identify genuine Nile tilapia body features which differ in dependence of the animals habitat. We find that automatically inferred Nile tilapia body features corroborate and expand the results of a landmark based analysis that the anterior dorsum, the fish belly, the posterior dorsal region and the caudal fin show signs of adaptation to the fish habitat. We may thus conclude that Nile tilapia show habitat specific morphotypes and that a ML analysis allows inferring novel biological knowledge in a reproducible manner.</p>]]></description>
            <pubDate><![CDATA[2021-04-15T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Enhanced streamflow prediction with SWAT using support vector regression for spatial calibration: A case study in the Illinois River watershed, U.S.]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766030902241-cac7fafd-0c16-4691-9c81-26bd9d085fca/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0248489</link>
            <description><![CDATA[<p class="para" id="N65539">Accurate streamflow prediction plays a pivotal role in hydraulic project design, nonpoint source pollution estimation, and water resources planning and management. However, the highly non-linear relationship between rainfall and runoff makes prediction difficult with desirable accuracy. To improve the accuracy of monthly streamflow prediction, a seasonal Support Vector Regression (SVR) model coupled to the Soil and Water Assessment Tool (SWAT) model was developed for 13 subwatersheds in the Illinois River watershed (IRW), U.S. Terrain, precipitation, soil, land use and land cover, and monthly streamflow data were used to build the SWAT model. SWAT Streamflow output and the upstream drainage area were used as two input variables into SVR to build the hybrid SWAT-SVR model. The Calibration Uncertainty Procedure (SWAT-CUP) and Sequential Uncertainty Fitting-2 (SUFI-2) algorithms were applied to compare the model performance against SWAT-SVR. The spatial calibration and leave-one-out sampling methods were used to calibrate and validate the hybrid SWAT-SVR model. The results showed that the SWAT-SVR model had less deviation and better performance than SWAT-CUP simulations. SWAT-SVR predicted streamflow more accurately during the wet season than the dry season. The model worked well when it was applied to simulate medium flows with discharge between 5 m<sup>3</sup> s<sup>-1</sup> and 30 m<sup>3</sup> s<sup>-1</sup>, and its applicable spatial scale fell between 500 to 3000 km<sup>2</sup>. The overall performance of the model on yearly time series is “Satisfactory”. This new SWAT-SVR model has not only the ability to capture intrinsic non-linear behaviors between rainfall and runoff while considering the mechanism of runoff generation but also can serve as a reliable regional tool for an ungauged or limited data watershed that has similar hydrologic characteristics with the IRW.</p>]]></description>
            <pubDate><![CDATA[2021-04-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[A direct comparison of theory-driven and machine learning prediction of suicide: A meta-analysis]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766030840151-82ff96fd-336f-40ac-b0ca-9958619fb8d4/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249833</link>
            <description><![CDATA[<p class="para" id="N65539">Theoretically-driven models of suicide have long guided suicidology; however, an approach employing machine learning models has recently emerged in the field. Some have suggested that machine learning models yield improved prediction as compared to theoretical approaches, but to date, this has not been investigated in a systematic manner. The present work directly compares widely researched theories of suicide (<i>i</i>.<i>e</i>., BioSocial, Biological, Ideation-to-Action, and Hopelessness Theories) to machine learning models, comparing the accuracy between the two differing approaches. We conducted literature searches using PubMed, PsycINFO, and Google Scholar, gathering effect sizes from theoretically-relevant constructs and machine learning models. Eligible studies were longitudinal research articles that predicted suicide ideation, attempts, or death published prior to May 1, 2020. 124 studies met inclusion criteria, corresponding to 330 effect sizes. Theoretically-driven models demonstrated suboptimal prediction of ideation (wOR = 2.87; 95% CI, 2.65–3.09; <i>k</i> = 87), attempts (wOR = 1.43; 95% CI, 1.34–1.51; <i>k</i> = 98), and death (wOR = 1.08; 95% CI, 1.01–1.15; <i>k</i> = 78). Generally, Ideation-to-Action (<i>w</i>OR = 2.41, 95% CI = 2.21–2.64, <i>k</i> = 60) outperformed Hopelessness (<i>w</i>OR = 1.83, 95% CI 1.71–1.96, <i>k</i> = 98), Biological (<i>w</i>OR = 1.04; 95% CI .97–1.11, <i>k</i> = 100), and BioSocial (<i>w</i>OR = 1.32, 95% CI 1.11–1.58, <i>k</i> = 6) theories. Machine learning provided superior prediction of ideation (wOR = 13.84; 95% CI, 11.95–16.03; <i>k</i> = 33), attempts (wOR = 99.01; 95% CI, 68.10–142.54; <i>k</i> = 27), and death (wOR = 17.29; 95% CI, 12.85–23.27; <i>k</i> = 7). Findings from our study indicated that across all theoretically-driven models, prediction of suicide-related outcomes was suboptimal. Notably, among theories of suicide, theories within the Ideation-to-Action framework provided the most accurate prediction of suicide-related outcomes. When compared to theoretically-driven models, machine learning models provided superior prediction of suicide ideation, attempts, and death.</p>]]></description>
            <pubDate><![CDATA[2021-04-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Relay protection system of transmission line based on AI]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766023644771-81b30f55-5bfb-4085-a66c-acffab2c9a72/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246403</link>
            <description><![CDATA[<p class="para" id="N65539">With the development of modern power systems, higher requirements are imposed on relay protection technology. Traditional relay protection and fault diagnosis technologies have been unable to meet the requirements of the continuous development of power systems, and relay protection systems based on artificial intelligence(AI) technology have received increasing attention. Therefore, this document first analyses the weaknesses of traditional broadcast line protection and uses the adaptability and self-learning of artificial intelligence(AI); to propose the concept of protection of a relay line based on AI. In combination with the artificial nervous network, the AI-based relay protection system shall be studied and the experimental model shall be developed. This paper validates it with simulation experiments. The research results show that for the analysis of the ANN test results of the subnetwork, the actual output of the subnetwork is very close to the ideal output, and the error does not exceed 0.2%. The system has good performance and high reliability.</p>]]></description>
            <pubDate><![CDATA[2021-04-07T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Modeling the spatial distribution of anthrax in southern Kenya]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766023402025-ee63126c-d7f6-48ad-9b76-99778918b368/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pntd.0009301</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Background</h3><p class="para" id="N65543">Anthrax is an important zoonotic disease in Kenya associated with high animal and public health burden and widespread socio-economic impacts. The disease occurs in sporadic outbreaks that involve livestock, wildlife, and humans, but knowledge on factors that affect the geographic distribution of these outbreaks is limited, challenging public health intervention planning.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Methods</h3><p class="para" id="N65549">Anthrax surveillance data reported in southern Kenya from 2011 to 2017 were modeled using a boosted regression trees (BRT) framework. An ensemble of 100 BRT experiments was developed using a variable set of 18 environmental covariates and 69 unique anthrax locations. Model performance was evaluated using AUC (area under the curve) ROC (receiver operating characteristics) curves.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Results</h3><p class="para" id="N65555">Cattle density, rainfall of wettest month, soil clay content, soil pH, soil organic carbon, length of longest dry season, vegetation index, temperature seasonality, in order, were identified as key variables for predicting environmental suitability for anthrax in the region. BRTs performed well with a mean AUC of 0.8. Areas highly suitable for anthrax were predicted predominantly in the southwestern region around the shared Kenya-Tanzania border and a belt through the regions and highlands in central Kenya. These suitable regions extend westwards to cover large areas in western highlands and the western regions around Lake Victoria and bordering Uganda. The entire eastern and lower-eastern regions towards the coastal region were predicted to have lower suitability for anthrax.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Conclusion</h3><p class="para" id="N65561">These modeling efforts identified areas of anthrax suitability across southern Kenya, including high and medium agricultural potential regions and wildlife parks, important for tourism and foreign exchange. These predictions are useful for policy makers in designing targeted surveillance and/or control interventions in Kenya.</p><p class="para" id="N65563">We thank the staff of Directorate of Veterinary Services under the Ministry of Agriculture, Livestock and Fisheries, for collecting and providing the anthrax historical occurrence data.</p></div><p class="para" id="N65542">Anthrax is a neglected zoonosis worldwide. In Kenya, outbreaks have been reported in wildlife, livestock, and humans, resulting in severe public health burden and socio-economic impacts. Because of this, anthrax is ranked as the highest priority disease in the country. To identify factors that influence the spatial distribution of the disease in Kenya, we analyzed surveillance on available anthrax outbreaks recorded in the southern half of the country. Areas predicted to be highly suitable for the disease were predominantly in the southwestern region around the shared Kenya-Tanzania border running as a belt through central regions and central highlands of Kenya. These suitability regions extend westwards to cover large areas in western highlands and the western regions around Lake Victoria and bordering Uganda. The entire eastern and lower-eastern regions towards the coastal region were predicted to have lower suitability for anthrax. Cattle density, rainfall of wettest month, soil clay content, soil pH, soil organic carbon, length of longest dry season, vegetation index and temperature seasonality were key variables predicting the distribution of anthrax in the region. The study generated a suitability map depicting geographical areas that can be targeted for risk-based surveillance and or control measures for the disease.</p>]]></description>
            <pubDate><![CDATA[2021-03-29T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Pandemic velocity: Forecasting COVID-19 in the US with a machine learning &amp; Bayesian time series compartmental model]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766022366991-200f7bab-04a4-4807-b88a-8dce7f793684/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pcbi.1008837</link>
            <description><![CDATA[<p class="para" id="N65539">Predictions of COVID-19 case growth and mortality are critical to the decisions of political leaders, businesses, and individuals grappling with the pandemic. This predictive task is challenging due to the novelty of the virus, limited data, and dynamic political and societal responses. We embed a Bayesian time series model and a random forest algorithm within an epidemiological compartmental model for empirically grounded COVID-19 predictions. The Bayesian case model fits a location-specific curve to the velocity (first derivative) of the log transformed cumulative case count, borrowing strength across geographic locations and incorporating prior information to obtain a posterior distribution for case trajectories. The compartmental model uses this distribution and predicts deaths using a random forest algorithm trained on COVID-19 data and population-level characteristics, yielding daily projections and interval estimates for cases and deaths in U.S. states. We evaluated the model by training it on progressively longer periods of the pandemic and computing its predictive accuracy over 21-day forecasts. The substantial variation in predicted trajectories and associated uncertainty between states is illustrated by comparing three unique locations: New York, Colorado, and West Virginia. The sophistication and accuracy of this COVID-19 model offer reliable predictions and uncertainty estimates for the current trajectory of the pandemic in the U.S. and provide a platform for future predictions as shifting political and societal responses alter its course.</p><p class="para" id="N65542">COVID-19 models can be roughly classified as mathematical models that simulate disease within a population, including epidemiological compartmental models, or statistical curve-fitting models that fit a function to observed data and extrapolate forward into the future. Bridging this divide, we combine the strengths of curve-fitting statistical models and the structure of epidemiological models, by embedding a Bayesian velocity model and a machine learning algorithm (random forest) into the framework of a compartmental model. Fusing these models together exploits the particular strengths of each to glean as much information as possible from the currently available data. We identify the velocity of log cumulative cases as an excellent target for modeling and extrapolating COVID-19 case trajectories. We empirically evaluate the predictive performance of the model and provide predicted trajectories with credible intervals for cumulative confirmed case count, active confirmed infections and COVID-19 deaths for each of the 50 U.S. states. Combining sophisticated data analytic methods with proven epidemiological models offers an empirically grounded strategy for making realistic predictions and quantifying their uncertainty. These predictions indicate substantial variation in the COVID-19 trajectories of U.S. states.</p>]]></description>
            <pubDate><![CDATA[2021-03-29T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Using the antibody-antigen binding interface to train image-based deep neural networks for antibody-epitope classification]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766022123607-edeacd5a-cc6e-4f07-a933-410e0f81ec2b/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pcbi.1008864</link>
            <description><![CDATA[<p class="para" id="N65539">High-throughput B-cell sequencing has opened up new avenues for investigating complex mechanisms underlying our adaptive immune response. These technological advances drive data generation and the need to mine and analyze the information contained in these large datasets, in particular the identification of therapeutic antibodies (Abs) or those associated with disease exposure and protection. Here, we describe our efforts to use artificial intelligence (AI)-based image-analyses for prospective classification of Abs based solely on sequence information. We hypothesized that Abs recognizing the same part of an antigen share a limited set of features at the binding interface, and that the binding site regions of these Abs share share common structure and physicochemical property patterns that can serve as a “fingerprint” to recognize uncharacterized Abs. We combined large-scale sequence-based protein-structure predictions to generate ensembles of 3-D Ab models, reduced the Ab binding interface to a 2-D image (fingerprint), used pre-trained convolutional neural networks to extract features, and trained deep neural networks (DNNs) to classify Abs. We evaluated this approach using Ab sequences derived from human HIV and Ebola viral infections to differentiate between two Abs, Abs belonging to specific B-cell family lineages, and Abs with different epitope preferences. In addition, we explored a different type of DNN method to detect one class of Abs from a larger pool of Abs. Testing on Ab sets that had been kept aside during model training, we achieved average prediction accuracies ranging from 71–96% depending on the complexity of the classification task. The high level of accuracies reached during these classification tests suggests that the DNN models were able to learn a series of structural patterns shared by Abs belonging to the same class. The developed methodology provides a means to apply AI-based image recognition techniques to analyze high-throughput B-cell sequencing datasets (repertoires) for Ab classification.</p><p class="para" id="N65542">The ability to take advantage of the rapid progress in AI for biological and medical application oftentimes requires looking at the problem from a non-traditional point-of-view. The adaptive immune system plays a key role in providing long-term immunity against pathogens. The repertoire of circulating B-cells that produce unique pathogen-specific antibodies in an individual contains immense information on both the status of the immune response at particular time and that individual’s immune history. With high-throughput sequencing, we can now obtain Ab sequences for thousands of B cells from a single patient blood sample, but functionally characterizing antibodies on this scale remains on daunting task. Here, we propose to use AI to functionally classify Abs from sequence alone by re-casting this classification problem as an image recognition problem. Just as traditional image recognition involves training AI to distinguish different types of objects, we sought to use AI to distinguish different types of Ab-antigen binding interfaces. Towards that end, we generated ensembles of Ab structures from sequence, and generated 2-D ‘fingerprints’ of each structure that captures the essential molecular and chemical structure of the Ab binding site regions, and trained a Convolution and Deep Neural Network based AI model to classify Ab fingerprints associated with different functional characteristics. We applied this DNN-based approach to accurately predict antibody family lineage and epitope specificity against Ebola and HIV-1 viruses, and to detect sequence-diverse antibodies with similar binding properties as the ones we used for training.</p>]]></description>
            <pubDate><![CDATA[2021-03-29T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Development of a graph convolutional neural network model for efficient prediction of protein-ligand binding affinities]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766009944995-db8eaf5d-56cc-403f-8fd3-7c84243aa7fc/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249404</link>
            <description><![CDATA[<p class="para" id="N65539">Prediction of protein-ligand interactions is a critical step during the initial phase of drug discovery. We propose a novel deep-learning-based prediction model based on a graph convolutional neural network, named GraphBAR, for protein-ligand binding affinity. Graph convolutional neural networks reduce the computational time and resources that are normally required by the traditional convolutional neural network models. In this technique, the structure of a protein-ligand complex is represented as a graph of multiple adjacency matrices whose entries are affected by distances, and a feature matrix that describes the molecular properties of the atoms. We evaluated the predictive power of GraphBAR for protein-ligand binding affinities by using PDBbind datasets and proved the efficiency of the graph convolution. Given the computational efficiency of graph convolutional neural networks, we also performed data augmentation to improve the model performance. We found that data augmentation with docking simulation data could improve the prediction accuracy although the improvement seems not to be significant. The high prediction performance and speed of GraphBAR suggest that such networks can serve as valuable tools in drug discovery.</p>]]></description>
            <pubDate><![CDATA[2021-04-08T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[A computational lens into how music characterizes genre in film]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766009409773-956cb068-6f65-45b3-ab3e-801ddaa34773/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249957</link>
            <description><![CDATA[<p class="para" id="N65539">Film music varies tremendously across genre in order to bring about different responses in an audience. For instance, composers may evoke passion in a romantic scene with lush string passages or inspire fear throughout horror films with inharmonious drones. This study investigates such phenomena through a quantitative evaluation of music that is associated with different film genres. We construct supervised neural network models with various pooling mechanisms to predict a film’s genre from its soundtrack. We use these models to compare handcrafted music information retrieval (MIR) features against VGGish audio embedding features, finding similar performance with the top-performing architectures. We examine the best-performing MIR feature model through permutation feature importance (PFI), determining that mel-frequency cepstral coefficient (MFCC) and tonal features are most indicative of musical differences between genres. We investigate the interaction between musical and visual features with a cross-modal analysis, and do not find compelling evidence that music characteristic of a certain genre implies low-level visual features associated with that genre. Furthermore, we provide software code to replicate this study at https://github.com/usc-sail/mica-music-in-media. This work adds to our understanding of music’s use in multi-modal contexts and offers the potential for future inquiry into human affective experiences.</p>]]></description>
            <pubDate><![CDATA[2021-04-08T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Predicting antimicrobial mechanism-of-action from transcriptomes: A generalizable explainable artificial intelligence approach]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766009185078-d7853285-378a-4c63-b0ee-53f974c5129a/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pcbi.1008857</link>
            <description><![CDATA[<p class="para" id="N65539">To better combat the expansion of antibiotic resistance in pathogens, new compounds, particularly those with novel mechanisms-of-action [MOA], represent a major research priority in biomedical science. However, rediscovery of known antibiotics demonstrates a need for approaches that accurately identify potential novelty with higher throughput and reduced labor. Here we describe an explainable artificial intelligence classification methodology that emphasizes prediction performance and human interpretability by using a Hierarchical Ensemble of Classifiers model optimized with a novel feature selection algorithm called <i>Clairvoyance</i>; collectively referred to as a CoHEC model. We evaluated our methods using whole transcriptome responses from <i>Escherichia coli</i> challenged with 41 known antibiotics and 9 crude extracts while depositing 122 transcriptomes unique to this study. Our CoHEC model can properly predict the primary MOA of previously unobserved compounds in both purified forms and crude extracts at an accuracy above 99%, while also correctly identifying darobactin, a newly discovered antibiotic, as having a novel MOA. In addition, we deploy our methods on a recent <i>E</i>. <i>coli</i> transcriptomics dataset from a different strain and a <i>Mycobacterium smegmatis</i> metabolomics timeseries dataset showcasing exceptionally high performance; improving upon the performance metrics of the original publications. We not only provide insight into the biological interpretation of our model but also that the concept of MOA is a non-discrete heuristic with diverse effects for different compounds within the same MOA, suggesting substantial antibiotic diversity awaiting discovery within existing MOA.</p><p class="para" id="N65542">As antimicrobial resistance is on the rise, the need for compounds with novel targets or mechanisms-of-action [MOA] are of the utmost importance from the standpoint of public health. A major bottleneck in drug discovery is the ability to rapidly screen candidate compounds for precise MOA activity as current approaches are expensive, time consuming, and are difficult to implement in high-throughput. To alleviate this bottleneck in drug discovery, we developed a human interpretable artificial intelligence classification framework that can be used to build highly accurate and flexible predictive models. In this study, we investigated antimicrobial MOA through the transcriptional responses of <i>Escherichia coli</i> challenged with 41 known antibiotic compounds, 9 crude extracts, and a recently discovered (circa 2019) compound, darobactin, with novel MOA activity. We implemented a highly stringent Leave Compound Out Cross-Validation procedure to stress-test our predictive models by simulating the scenario of observing novel compounds. Furthermore, we developed a versatile feature selection algorithm, <i>Clairvoyance</i>, that we apply to our hierarchical ensemble of classifiers framework to build high performance explainable machine-learning models. Although the methods in this study were developed and stress-tested to predict the primary MOA from transcriptomic responses in <i>E</i>. <i>coli</i>, we designed these methods for general application to any classification problem and open-sourced the implementations in our <i>Soothsayer</i> Python package. We further demonstrate the versatility of these methods by deploying them on recent <i>Mycobacterium smegmatis</i> metabolomic and <i>E</i>. <i>coli</i> transcriptomics datasets to predict MOA with high accuracy.</p>]]></description>
            <pubDate><![CDATA[2021-03-29T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[RBF neural network based backstepping terminal sliding mode MPPT control technique for PV system]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766009160747-618a69f2-b757-4e4e-a974-fb443b2e0200/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249705</link>
            <description><![CDATA[<p class="para" id="N65539">The energy demand in the world has increased rapidly in the last few decades. This demand is arising the need for alternative energy resources. Solar energy is the most eminent energy resource which is completely free from pollution and fuel. However, the problem occurs when it comes to efficiency under different atmospheric conditions such as varying temperature and solar irradiance. To achieve its maximum efficiency, an algorithm of maximum power point tracking (MPPT) is needed to fetch maximum power from the photovoltaic (PV) system. In this article, a nonlinear backstepping terminal sliding mode control (BTSMC) is proposed for maximum power extraction. The system is finite-time stable and its stability is validated through the Lyapunov function. A DC-DC buck-boost converter is used to deliver PV power to the load. For the proposed controller, reference voltages are generated by a radial basis function neural network (RBF NN). The proposed controller performance is tested using the MATLAB/Simulink tool. Furthermore, the controller performance is compared with the perturb and observe (P&amp;O) MPPT algorithm, Proportional Integral Derivative (PID) controller and backstepping MPPT nonlinear controller. The results validate that the proposed controller offers better tracking and fast convergence in finite time under rapidly varying conditions of the environment.</p>]]></description>
            <pubDate><![CDATA[2021-04-08T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[A simple interpretation of undirected edges in essential graphs is wrong]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766008496780-31564b6d-4b98-4556-8c25-ac25db4769c0/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249415</link>
            <description><![CDATA[<p class="para" id="N65539">Artificial intelligence for causal discovery frequently uses Markov equivalence classes of directed acyclic graphs, graphically represented as <i>essential graphs</i>, as a way of representing uncertainty in causal directionality. There has been confusion regarding how to interpret undirected edges in essential graphs, however. In particular, experts and non-experts both have difficulty quantifying the likelihood of uncertain causal arrows being pointed in one direction or another. A simple interpretation of undirected edges treats them as having equal odds of being oriented in either direction, but I show in this paper that any agent interpreting undirected edges in this simple way can be Dutch booked. In other words, I can construct a set of bets that appears rational for the users of the simple interpretation to accept, but for which in all possible outcomes they lose money. I put forward another interpretation, prove this interpretation leads to a bet-taking strategy that is sufficient to avoid all Dutch books of this kind, and conjecture that this strategy is also necessary for avoiding such Dutch books. Finally, I demonstrate that undirected edges that are more likely to be oriented in one direction than the other are common in graphs with 4 nodes and 3 edges.</p>]]></description>
            <pubDate><![CDATA[2021-04-08T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[A metric learning method for estimating myelin content based on T2-weighted MRI from a de- and re-myelination model of multiple sclerosis]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766005883519-17994936-9e4d-4090-beef-0ec8db622a13/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249460</link>
            <description><![CDATA[<p class="para" id="N65539">Myelin plays a critical role in the pathogenesis of neurological disorders but is difficult to characterize in vivo using standard analysis methods. Our goal was to develop a novel analytical framework for estimating myelin content using T2-weighted magnetic resonance imaging (MRI) based on a de- and re-myelination model of multiple sclerosis. We examined 18 mice with lysolecithin induced demyelination and spontaneous remyelination in the ventral white matter of thoracic spinal cord. Cohorts of 6 mice underwent 9.4T MRI at days 7 (peak demyelination), 14 (ongoing recovery), and 28 (near complete recovery), as well as histological analysis of myelin and the associated cellularity at corresponding timepoints. Our MRI framework took an unsupervised learning approach, including tissue segmentation using a Gaussian Markov random field (GMRF), and myelin and cellularity feature estimation based on the Mahalanobis distance. For comparison, we also investigated 2 regression-based supervised learning approaches, one using our GMRF results, and another using a freely available generalized additive model (GAM). Results showed that GMRF segmentation was 73.2% accurate, and our unsupervised learning method achieved a correlation coefficient of 0.67 (top quartile: 0.78) with histological myelin, similar to 0.70 (top quartile: 0.78) obtained using supervised analyses. Further, the area under the receiver operator characteristic curve of our unsupervised myelin feature (0.883, 95% CI: 0.874–0.891) was significantly better than any of the supervised models in detecting white matter myelin as compared to histology. Collectively, metric learning using standard MRI may prove to be a new alternative method for estimating myelin content, which ultimately can improve our disease monitoring ability in a clinical setting.</p>]]></description>
            <pubDate><![CDATA[2021-04-05T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Deep learning classification of lipid droplets in quantitative phase images]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766005839581-76f5077c-9f7b-401c-9395-1e55a4702933/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249196</link>
            <description><![CDATA[<p class="para" id="N65539">We report the application of supervised machine learning to the automated classification of lipid droplets in label-free, quantitative-phase images. By comparing various machine learning methods commonly used in biomedical imaging and remote sensing, we found convolutional neural networks to outperform others, both quantitatively and qualitatively. We describe our imaging approach, all implemented machine learning methods, and their performance with respect to computational efficiency, required training resources, and relative method performance measured across multiple metrics. Overall, our results indicate that quantitative-phase imaging coupled to machine learning enables accurate lipid droplet classification in single living cells. As such, the present paradigm presents an excellent alternative of the more common fluorescent and Raman imaging modalities by enabling label-free, ultra-low phototoxicity, and deeper insight into the thermodynamics of metabolism of single cells.</p>]]></description>
            <pubDate><![CDATA[2021-04-05T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Innovation indicators based on firm websites—Which website characteristics predict firm-level innovation activity?]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766005822007-aa86d107-46ee-40c7-a67d-36b3533cdbed/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249583</link>
            <description><![CDATA[<p class="para" id="N65539">Web-based innovation indicators may provide new insights into firm-level innovation activities. However, little is known yet about the accuracy and relevance of web-based information for measuring innovation. In this study, we use data on 4,487 firms from the Mannheim Innovation Panel (MIP) 2019, the German contribution to the European Community Innovation Survey (CIS), to analyze which website characteristics perform as predictors of innovation activity at the firm level. Website characteristics are measured by several data mining methods and are used as features in different Random Forest classification models that are compared against each other. Our results show that the most relevant website characteristics are textual content, the use of English language, the number of subpages and the amount of characters on a website. In our main analysis, models using all website characteristics jointly yield AUC values of up to 0.75 and increase accuracy scores by up to 18 percentage points compared to a baseline prediction based on the sample mean. Moreover, predictions with website characteristics significantly differ from baseline predictions according to a McNemar test. Results also indicate a better performance for the prediction of product innovators and firms with innovation expenditures than for the prediction of process innovators.</p>]]></description>
            <pubDate><![CDATA[2021-04-05T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Investigation of ANN architecture for predicting shear strength of fiber reinforcement bars concrete beams]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766003645738-6caf71f6-a06f-4f7b-b0ac-d79420fe2844/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247391</link>
            <description><![CDATA[<p class="para" id="N65539">In this paper, an extensive simulation program is conducted to find out the optimal ANN model to predict the shear strength of fiber-reinforced polymer (FRP) concrete beams containing both flexural and shear reinforcements. For acquiring this purpose, an experimental database containing 125 samples is collected from the literature and used to find the best architecture of ANN. In this database, the input variables consist of 9 inputs, such as the ratio of the beam width, the effective depth, the shear span to the effective depth, the compressive strength of concrete, the longitudinal FRP reinforcement ratio, the modulus of elasticity of longitudinal FRP reinforcement, the FRP shear reinforcement ratio, the tensile strength of FRP shear reinforcement, the modulus of elasticity of FRP shear reinforcement. Thereafter, the selection of the appropriate architecture of ANN model is performed and evaluated by common statistical measurements. The results show that the optimal ANN model is a highly efficient predictor of the shear strength of FRP concrete beams with a maximum R<sup>2</sup> value of 0.9634 on the training part and an R<sup>2</sup> of 0.9577 on the testing part, using the best architecture. In addition, a sensitivity analysis using the optimal ANN model over 500 Monte Carlo simulations is performed to interpret the influence of reinforcement type on the stability and accuracy of ANN model in predicting shear strength. The results of this investigation could facilitate and enhance the use of ANN model in different real-world problems in the field of civil engineering.</p>]]></description>
            <pubDate><![CDATA[2021-04-02T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Predicting student satisfaction of emergency remote learning in higher education during COVID-19 using machine learning techniques]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766003362922-353ea0ed-5486-49ba-aeea-00b31b526097/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0249423</link>
            <description><![CDATA[<p class="para" id="N65539">Despite the wide adoption of emergency remote learning (ERL) in higher education during the COVID-19 pandemic, there is insufficient understanding of influencing factors predicting student satisfaction for this novel learning environment in crisis. The present study investigated important predictors in determining the satisfaction of undergraduate students (N = 425) from multiple departments in using ERL at a self-funded university in Hong Kong while Moodle and Microsoft Team are the key learning tools. By comparing the predictive accuracy between multiple regression and machine learning models before and after the use of random forest recursive feature elimination, all multiple regression, and machine learning models showed improved accuracy while the most accurate model was the elastic net regression with 65.2% explained variance. The results show only neutral (4.11 on a 7-point Likert scale) regarding the overall satisfaction score on ERL. Even majority of students are competent in technology and have no obvious issue in accessing learning devices or Wi-Fi, face-to-face learning is more preferable compared to ERL and this is found to be the most important predictor. Besides, the level of efforts made by instructors, the agreement on the appropriateness of the adjusted assessment methods, and the perception of online learning being well delivered are shown to be highly important in determining the satisfaction scores. The results suggest that the need of reviewing the quality and quantity of modified assessment accommodated for ERL and structured class delivery with the suitable amount of interactive learning according to the learning culture and program nature.</p>]]></description>
            <pubDate><![CDATA[2021-04-02T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Application of an emotional classification model in e-commerce text based on an improved transformer model]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1766000173130-c8327e7f-c9e5-4e96-bfeb-c2cff3b79af8/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247984</link>
            <description><![CDATA[<p class="para" id="N65539">With the rapid development of the mobile internet, people are becoming more dependent on the internet to express their comments on products or stores; meanwhile, text sentiment classification of these comments has become a research hotspot. In existing methods, it is fairly popular to apply a deep learning method to the text classification task. Aiming at solving information loss, weak context and other problems, this paper makes an improvement based on the transformer model to reduce the difficulty of model training and training time cost and achieve higher overall model recall and accuracy in text sentiment classification. The transformer model replaces the traditional convolutional neural network (CNN) and the recurrent neural network (RNN) and is fully based on the attention mechanism; therefore, the transformer model effectively improves the training speed and reduces training difficulty. This paper selects e-commerce reviews as research objects and applies deep learning theory. First, the text is preprocessed by word vectorization. Then the IN standardized method and the GELUs activation function are applied based on the original model to analyze the emotional tendencies of online users towards stores or products. The experimental results show that our method improves by 9.71%, 6.05%, 5.58% and 5.12% in terms of recall and approaches the peak level of the F1 value in the test model by comparing BiLSTM, Naive Bayesian Model, the serial BiLSTM_CNN model and BiLSTM with an attention mechanism model. Therefore, this finding proves that our method can be used to improve the text sentiment classification accuracy and effectively apply the method to text classification.</p>]]></description>
            <pubDate><![CDATA[2021-03-05T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Using machine learning to improve risk prediction in durable left ventricular assist devices]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765999737724-97e68969-037c-492a-94b0-3404044b7b16/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247866</link>
            <description><![CDATA[<p class="para" id="N65539">Risk models have historically displayed only moderate predictive performance in estimating mortality risk in left ventricular assist device therapy. This study evaluated whether machine learning can improve risk prediction for left ventricular assist devices. Primary durable left ventricular assist devices reported in the Interagency Registry for Mechanically Assisted Circulatory Support between March 1, 2006 and December 31, 2016 were included. The study cohort was randomly divided 3:1 into training and testing sets. Logistic regression and machine learning models (extreme gradient boosting) were created in the training set for 90-day and 1-year mortality and their performance was evaluated after bootstrapping with 1000 replications in the testing set. Differences in model performance were also evaluated in cases of concordance versus discordance in predicted risk between logistic regression and extreme gradient boosting as defined by equal size patient tertiles. A total of 16,120 patients were included. Calibration metrics were comparable between logistic regression and extreme gradient boosting. C-index was improved with extreme gradient boosting (90-day: 0.707 [0.683–0.730] versus 0.740 [0.717–0.762] and 1-year: 0.691 [0.673–0.710] versus 0.714 [0.695–0.734]; each p&lt;0.001). Net reclassification index analysis similarly demonstrated an improvement of 48.8% and 36.9% for 90-day and 1-year mortality, respectively, with extreme gradient boosting (each p&lt;0.001). Concordance in predicted risk between logistic regression and extreme gradient boosting resulted in substantially improved c-index for both logistic regression and extreme gradient boosting (90-day logistic regression 0.536 versus 0.752, 1-year logistic regression 0.555 versus 0.726, 90-day extreme gradient boosting 0.623 versus 0.772, 1-year extreme gradient boosting 0.613 versus 0.742, each p&lt;0.001). These results demonstrate that machine learning can improve risk model performance for durable left ventricular assist devices, both independently and as an adjunct to logistic regression.</p>]]></description>
            <pubDate><![CDATA[2021-03-10T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Probabilistic social learning improves the public’s judgments of news veracity]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765999374024-987afeef-5578-4956-adf1-32d309755225/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247487</link>
            <description><![CDATA[<p class="para" id="N65539">The digital spread of misinformation is one of the leading threats to democracy, public health, and the global economy. Popular strategies for mitigating misinformation include crowdsourcing, machine learning, and media literacy programs that require social media users to classify news in binary terms as either true or false. However, research on peer influence suggests that framing decisions in binary terms can amplify judgment errors and limit social learning, whereas framing decisions in probabilistic terms can reliably improve judgments. In this preregistered experiment, we compare online peer networks that collaboratively evaluated the veracity of news by communicating either binary or probabilistic judgments. Exchanging probabilistic estimates of news veracity substantially improved individual and group judgments, with the effect of eliminating polarization in news evaluation. By contrast, exchanging binary classifications reduced social learning and maintained polarization. The benefits of probabilistic social learning are robust to participants’ education, gender, race, income, religion, and partisanship.</p>]]></description>
            <pubDate><![CDATA[2021-03-09T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Application of deep learning in automatic detection of technical and tactical indicators of table tennis]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765999058388-ca7b58e1-c9fd-4826-b8c0-ce96c5760d5b/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245259</link>
            <description><![CDATA[<p class="para" id="N65539">A DCNN-LSTM (Deep Convolutional Neural Network-Long Short Term Memory) model is proposed to recognize and track table tennis’s real-time trajectory in complex environments, aiming to help the audiences understand competition details and provide a reference for training enthusiasts using computers. Real-time motion features are extracted via deep reinforcement networks. DCNN tracks the recognized objects, and the LSTM algorithm predicts the ball’s trajectory. The model is tested on a self-built video dataset and existing systems and compared with other algorithms to verify its effectiveness. Finally, an overall tactical detection system is built to measure ball rotation and predict ball trajectory. Results demonstrate that in feature extraction, the Deep Deterministic Policy Gradient (DDPG) algorithm has the best performance, with a maximum accuracy rate of 89% and a minimum mean square error of 0.2475. The accuracy of target tracking effect and trajectory prediction is as high as 90%. Compared with traditional methods, the performance of the DCNN-LSTM model based on deep learning is improved by 23.17%. The implemented automatic detection system of table tennis tactical indicators can deal with the problems of table tennis tracking and rotation measurement. It can provide a theoretical foundation and practical value for related research in real-time dynamic detection of balls.</p>]]></description>
            <pubDate><![CDATA[2021-03-09T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The adoption of cryptocurrency as a disruptive force: Deep learning-based dual stage structural equation modelling and artificial neural network analysis]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765998584521-48d32a10-bceb-42e2-990a-1fa12cc4d971/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247582</link>
            <description><![CDATA[<p class="para" id="N65539">In recent years, the growth of cryptocurrency has undergone an enormous increase in cryptocurrency markets all around the world. Sadly, only insignificant heed has been paid to the unveiling of determinants of cryptocurrency adoption globally, particularly in emerging markets like Malaysia. The purpose of the study is to examine whether the application of deep learning-based dual-stage Partial Least Square-Structural Equation Modelling (PLS-SEM) &amp; Artificial Neural Network (ANN) analysis enable better in-depth research results as compared to single-step PLS-SEM approach and to excavate factors which can predict behavioural intention to adopt cryptocurrency. The Unified Theory of Acceptance and Use of Technology 2 (UTAUT2) model were extended with the inclusion of trust and personnel innovativeness. The model was further validated by introducing a new path model compared to the original UTAUT2 model and the moderating role of personal innovativeness between performance expectancy and price value, with a sample of 314 respondents. Contrary to previous technology adoption studies that used PLS-SEM &amp; ANN as single-stage analysis, this study further enhanced the analysis by applying a deep learning-based dual-stage PLS-SEM and ANN method. The application of deep learning-based dual-stage PLS-SEM &amp; ANN analysis is a novel methodological approach, detecting both linear and non-linear associations among constructs. At the same time, it is regarded as a superior statistical approach as compared to traditional hybrid shallow SEM &amp; ANN single-stage analysis. Also, sensitivity analysis provides normalised importance using multi-layer perceptron with the feed-forward-back-propagation algorithm. Furthermore, the deep learning-based dual-stage PLS-SEM &amp; ANN revealed that trust proved to be the strongest predictor in driving user intention. The introduction of this new methodology and the theoretical contribution opens the vistas of the extant body of knowledge in technology-adoption related literature. This study also provides theoretical, practical and methodological contributions.</p>]]></description>
            <pubDate><![CDATA[2021-03-08T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[An apta-aggregation based machine learning assay for rapid quantification of lysozyme through texture parameters]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765995696104-e8032c36-1b07-46e0-8ac9-9f1cdb607fb8/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0248159</link>
            <description><![CDATA[<p class="para" id="N65539">A novel assay technique that involves quantification of lysozyme (Lys) through machine learning is put forward here. This article reports the tendency of the well- documented Ellington group anti-Lys aptamer, to produce aggregates when exposed to Lys. This property of apta-aggregation has been exploited here to develop an assay that quantifies the Lys using texture and area parameters from a photograph of the elliptical aggregate mass through machine learning. Two assay sets were made for the experimental procedure: one with high Lys concentration between 25–100 mM and another with low concentration between 1–20 mM. The high concentration set had a sample volume of 10 μl while the low concentration set had a higher sample volume of 100 μl, in order to obtain the statistical texture values reliably from the aggregate mass. The platform exhibited an experimental limit of detection of 1 mM and a response time of less than 10 seconds. Further, two potential operating modes for the aptamer were hypothesized for this aggregation property and the more accurate mode among the two was ascertained through bioinformatics studies.</p>]]></description>
            <pubDate><![CDATA[2021-03-08T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Attention based GRU-LSTM for software defect prediction]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765992272810-999de7bc-7f44-43fe-a48f-b2a631d10f51/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247444</link>
            <description><![CDATA[<p class="para" id="N65539">Software defect prediction (SDP) can be used to produce reliable, high-quality software. The current SDP is practiced on program granular components (such as file level, class level, or function level), which cannot accurately predict failures. To solve this problem, we propose a new framework called DP-AGL, which uses attention-based GRU-LSTM for statement-level defect prediction. By using clang to build an abstract syntax tree (AST), we define a set of 32 statement-level metrics. We label each statement, then make a three-dimensional vector and apply it as an automatic learning model, and then use a gated recurrent unit (GRU) with a long short-term memory (LSTM). In addition, the Attention mechanism is used to generate important features and improve accuracy. To verify our experiments, we selected 119,989 C/C++ programs in Code4Bench. The benchmark tests cover various programs and variant sets written by thousands of programmers. As an evaluation standard, compared with the state evaluation method, the recall, precision, accuracy and F1 measurement of our well-trained DP-AGL under normal conditions have increased by 1%, 4%, 5%, and 2% respectively.</p>]]></description>
            <pubDate><![CDATA[2021-03-04T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Prediction of femoral osteoporosis using machine-learning analysis with radiomics features and abdomen-pelvic CT: A retrospective single center preliminary study]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765991840000-eeee90c0-f681-48dc-bdbd-6d2d73a7ce82/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247330</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Background</h3><p class="para" id="N65543">Osteoporosis has increased and developed into a serious public health concern worldwide. Despite the high prevalence, osteoporosis is silent before major fragility fracture and the osteoporosis screening rate is low. Abdomen-pelvic CT (APCT) is one of the most widely conducted medical tests. Artificial intelligence and radiomics analysis have recently been spotlighted. This is the first study to evaluate the prediction performance of femoral osteoporosis using machine-learning analysis with radiomics features and APCT.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Materials and methods</h3><p class="para" id="N65549">500 patients (M: F = 70:430; mean age, 66.5 ± 11.8yrs; range, 50–96 years) underwent both dual-energy X-ray absorptiometry and APCT within 1 month. The volume of interest of the left proximal femur was extracted and 41 radiomics features were calculated using 3D volume of interest analysis. Top 10 importance radiomic features were selected by the intraclass correlation coefficient and random forest feature selection. Study cohort was randomly divided into 70% of the samples as the training cohort and the remaining 30% of the sample as the validation cohort. Prediction performance of machine-learning analysis was calculated using diagnostic test and comparison of area under the curve (AUC) of receiver operating characteristic curve analysis was performed between training and validation cohorts.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Results</h3><p class="para" id="N65555">The osteoporosis prevalence of this study cohort was 20.8%. The prediction performance of the machine-learning analysis to diagnose osteoporosis in the training and validation cohorts were as follows; accuracy, 92.9% vs. 92.7%; sensitivity, 86.6% vs. 80.0%; specificity, 94.5% vs. 95.8%; positive predictive value, 78.4% vs. 82.8%; and negative predictive value, 96.7% vs. 95.0%. The AUC to predict osteoporosis in the training and validation cohorts were 95.9% [95% confidence interval (CI), 93.7%-98.1%] and 96.0% [95% CI, 93.2%-98.8%], respectively, without significant differences (P = 0.962).</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Conclusion</h3><p class="para" id="N65561">Prediction performance of femoral osteoporosis using machine-learning analysis with radiomics features and APCT showed high validity with more than 93% accuracy, specificity, and negative predictive value.</p></div>]]></description>
            <pubDate><![CDATA[2021-03-04T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Hands-on training about overfitting]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765989232516-9e490f32-21aa-4740-a1fe-8429116970de/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pcbi.1008671</link>
            <description><![CDATA[<p class="para" id="N65539">Overfitting is one of the critical problems in developing models by machine learning. With machine learning becoming an essential technology in computational biology, we must include training about overfitting in all courses that introduce this technology to students and practitioners. We here propose a hands-on training for overfitting that is suitable for introductory level courses and can be carried out on its own or embedded within any data science course. We use workflow-based design of machine learning pipelines, experimentation-based teaching, and hands-on approach that focuses on concepts rather than underlying mathematics. We here detail the data analysis workflows we use in training and motivate them from the viewpoint of teaching goals. Our proposed approach relies on Orange, an open-source data science toolbox that combines data visualization and machine learning, and that is tailored for education in machine learning and explorative data analysis.</p><p class="para" id="N65542">Every teacher strives for an a-ha moment, a sudden revelation by the student who gained a fundamental insight she will always remember. In the past years, authors of this paper have been tailoring their courses in machine learning to include material that could lead students to such discoveries. We aim to expose machine learning to practitioners–not only computer scientists but also molecular biologists and students of biomedicine, that is, the end-users of bioinformatics’ computational approaches. In this article, we lay out a course that aims to teach about overfitting, one of the key concepts in machine learning that needs to be understood, mastered, and avoided in data science applications. We propose a hands-on approach that uses an open-source workflow-based data science toolbox that combines data visualization and machine learning. In the proposed training about overfitting, we first deceive the students, then expose the problem, and finally challenge them to find the solution. In the paper, we present three lessons in overfitting and associated data analysis workflows and motivate the use of introduced computation methods by relating them to concepts conveyed by instructors.</p>]]></description>
            <pubDate><![CDATA[2021-03-04T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Egg recognition: The importance of quantifying multiple repeatable features as visual identity signals]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765988481007-4fb59557-c961-474d-a032-4c6c0b90cf55/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0248021</link>
            <description><![CDATA[<p class="para" id="N65539">Brood parasitized and/or colonial birds use egg features as visual identity signals, which allow parents to recognize their own eggs and avoid paying fitness costs of misdirecting their care to others’ offspring. However, the mechanisms of egg recognition and discrimination are poorly understood. Most studies have put their focus on individual abilities to carry out these behavioural tasks, while less attention has been paid to the egg and how its signals may evolve to enhance its identification. We used 92 clutches (460 eggs) of the Eurasian coot <i>Fulica atra</i> to test whether eggs could be correctly classified into their corresponding clutches based only on their external appearance. Using SpotEgg, we characterized the eggs in 27 variables of colour, spottiness, shape and size from calibrated digital images. Then, we used these variables in a supervised machine learning algorithm for multi-class egg classification, where each egg was classified to the best matched clutch out of 92 studied clutches. The best model with all 27 explanatory variables assigned correctly 53.3% (CI = 42.6–63.7%) of eggs of the test-set, greatly exceeding the probability to classify the eggs by chance (1/92, 1.1%). This finding supports the hypothesis that eggs have visual identity signals in their phenotypes. Simplified models with fewer explanatory variables (10 or 15) showed lesser classification ability than full models, suggesting that birds may use multiple traits for egg recognition. Therefore, egg phenotypes should be assessed in their full complexity, including colour, patterning, shape and size. Most important variables for classification were those with the highest intraclutch correlation, demonstrating that individual recognition traits are repeatable. Algorithm classification performance improved by each extra training egg added to the model. Thus, repetition of egg design within a clutch would reinforce signals and would help females to create an internal template for true recognition of their own eggs. In conclusion, our novel approach based on machine learning provided important insights on how signallers broadcast their specific signature cues to enhance their recognisability.</p>]]></description>
            <pubDate><![CDATA[2021-03-04T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[A fault diagnosis method based on Auxiliary Classifier Generative Adversarial
Network for rolling bearing]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765983935551-3c82a526-f8f6-45d2-9762-6a23e3a63386/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246905</link>
            <description><![CDATA[<p class="para" id="N65539">Rolling bearing fault diagnosis is one of the challenging tasks and hot research topics
in the condition monitoring and fault diagnosis of rotating machinery. However, in
practical engineering applications, the working conditions of rotating machinery are
various, and it is difficult to extract the effective features of early fault due to the
vibration signal accompanied by high background noise pollution, and there are only a
small number of fault samples for fault diagnosis, which leads to the significant decline
of diagnostic performance. In order to solve above problems, by combining Auxiliary
Classifier Generative Adversarial Network (ACGAN) and Stacked Denoising Auto Encoder
(SDAE), a novel method is proposed for fault diagnosis. Among them, during the process of
training the ACGAN-SDAE, the generator and discriminator are alternately optimized through
the adversarial learning mechanism, which makes the model have significant diagnostic
accuracy and generalization ability. The experimental results show that our proposed
ACGAN-SDAE can maintain a high diagnosis accuracy under small fault samples, and have the
best adaptation performance across different load domains and better anti-noise
performance.</p>]]></description>
            <pubDate><![CDATA[2021-03-01T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Identification of success factors in elite wrestlers—An exploratory study]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765982948869-848dd171-9207-4f77-ba07-e7bf9073352a/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247565</link>
            <description><![CDATA[<p class="para" id="N65539">Identification of success factors in wrestling as well as establishing their hierarchy are crucial from a cognitive and practical standpoint. It may provide a lot of practical recommendations related to wrestling-specific training. The aim of this study was to identify and establish the hierarchy of success factors in wrestling regardless of a fighting style and weight class. This study included 168 elite male freestyle and Greco-Roman wrestlers. They were divided into two groups: athletes who won medals (successful wrestlers) in high-rank competitions (Polish Championships or higher) and those who did not win any medals (less successful wrestlers) in those competitions. The following elements were assessed: anthropological measurements, body composition, dynamic strength, strength endurance, agility, special endurance, wrestling-specific fitness, response time, technical wrestling skills and anaerobic capacity. For initial data analysis, one-way ANOVA (α = 0.005) was used. Random Forests classifier was employed to identify success factors and to determine the importance of each of these factors in terms of sports performance. Seven key success factors were identified: anaerobic power, strength endurance, response time, special endurance, wrestling-specific fitness and technical wrestling skills performed in a horizontal position. Random Forests turned out to be an effective method of modelling success in wrestling (compared to SVM and KNN, which were also used in the study). These findings suggest that wrestling-specific training can be effectively monitored by controlling several vital indicators of athletes’ preparedness: anaerobic power, strength endurance, response time, special endurance, wrestling-specific fitness and technical wrestling skills (the performance of reverse waistlock from a standing position and trunk grip gut wrench assessed by experts).</p>]]></description>
            <pubDate><![CDATA[2021-03-04T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Personalized prediction of early childhood asthma persistence: A machine learning approach]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765979339445-a7c95417-ef19-4dcd-afe6-c2da6235e7d5/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247784</link>
            <description><![CDATA[<p class="para" id="N65539">Early childhood asthma diagnosis is common; however, many children diagnosed before age 5 experience symptom resolution and it remains difficult to identify individuals whose symptoms will persist. Our objective was to develop machine learning models to identify which individuals diagnosed with asthma before age 5 continue to experience asthma-related visits. We curated a retrospective dataset for 9,934 children derived from electronic health record (EHR) data. We trained five machine learning models to differentiate individuals without subsequent asthma-related visits (transient diagnosis) from those with asthma-related visits between ages 5 and 10 (persistent diagnosis) given clinical information up to age 5 years. Based on average NPV-Specificity area (ANSA), all models performed significantly better than random chance, with XGBoost obtaining the best performance (0.43 mean ANSA). Feature importance analysis indicated age of last asthma diagnosis under 5 years, total number of asthma related visits, self-identified black race, allergic rhinitis, and eczema as important features. Although our models appear to perform well, a lack of prior models utilizing a large number of features to predict individual persistence makes direct comparison infeasible. However, feature importance analysis indicates our models are consistent with prior research indicating diagnosis age and prior health service utilization as important predictors of persistent asthma. We therefore find that machine learning models can predict which individuals will experience persistent asthma with good performance and may be useful to guide clinician and parental decisions regarding asthma counselling in early childhood.</p>]]></description>
            <pubDate><![CDATA[2021-03-01T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[A message-passing multi-task architecture for the implicit event and polarity detection]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765977616159-7976fa35-8af8-4d8c-a3b8-5a5d63782bed/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247704</link>
            <description><![CDATA[<p class="para" id="N65539">Implicit sentiment analysis is a challenging task because the sentiment of a text is expressed in a connotative manner. To tackle this problem, we propose to use textual events as a knowledge source to enrich network representations. To consider task interactions, we present a novel lightweight joint learning paradigm that can pass task-related messages between tasks during training iterations. This is distinct from previous methods that involve multi-task learning by simple parameter sharing. Besides, a human-annotated corpus with implicit sentiment labels and event labels is scarce, which hinders practical applications of deep neural models. Therefore, we further investigate a back-translation approach to expand training instances. Experiment results on a public benchmark demonstrate the effectiveness of both the proposed multi-task architecture and data augmentation strategy.</p>]]></description>
            <pubDate><![CDATA[2021-03-01T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[A natural language processing and deep learning approach to identify child abuse from pediatric electronic medical records]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765968911365-d7894a5a-fa66-4c37-aca1-b084f6de3fb5/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247404</link>
            <description><![CDATA[<p class="para" id="N65539">Child physical abuse is a leading cause of traumatic injury and death in children. In 2017, child abuse was responsible for 1688 fatalities in the United States, of 3.5 million children referred to Child Protection Services and 674,000 substantiated victims. While large referral hospitals maintain teams trained in Child Abuse Pediatrics, smaller community hospitals often do not have such dedicated resources to evaluate patients for potential abuse. Moreover, identification of abuse has a low margin of error, as false positive identifications lead to unwarranted separations, while false negatives allow dangerous situations to continue. This context makes the consistent detection of and response to abuse difficult, particularly given subtle signs in young, non-verbal patients. Here, we describe the development of artificial intelligence algorithms that use unstructured free-text in the electronic medical record—including notes from physicians, nurses, and social workers—to identify children who are suspected victims of physical abuse. Importantly, only the notes from time of first encounter (e.g.: birth, routine visit, sickness) to the last record before child protection team involvement were used. This allowed us to develop an algorithm using only information available prior to referral to the specialized child protection team. The study was performed in a multi-center referral pediatric hospital on patients screened for abuse within five different locations between 2015 and 2019. Of 1123 patients, 867 records were available after data cleaning and processing, and 55% were abuse-positive as determined by a multi-disciplinary team of clinical professionals. These electronic medical records were encoded with three natural language processing (NLP) algorithms—Bag of Words (BOW), Word Embeddings (WE), and Rules-Based (RB)—and used to train multiple neural network architectures. The BOW and WE encodings utilize the full free-text, while RB selects crucial phrases as identified by physicians. The best architecture was selected by average classification accuracy for the best performing model from each train-test split of a cross-validation experiment. Natural language processing coupled with neural networks detected cases of likely child abuse using only information available to clinicians prior to child protection team referral with average accuracy of 0.90±0.02 and average area under the receiver operator characteristic curve (ROC-AUC) 0.93±0.02 for the best performing Bag of Words models. The best performing rules-based models achieved average accuracy of 0.77±0.04 and average ROC-AUC 0.81±0.05, while a Word Embeddings strategy was severely limited by lack of representative embeddings. Importantly, the best performing model had a false positive rate of 8%, as compared to rates of 20% or higher in previously reported studies. This artificial intelligence approach can help screen patients for whom an abuse concern exists and streamline the identification of patients who may benefit from referral to a child protection team. Furthermore, this approach could be applied to develop computer-aided-diagnosis platforms for the challenging and often intractable problem of reliably identifying pediatric patients suffering from physical abuse.</p>]]></description>
            <pubDate><![CDATA[2021-02-26T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Spec2Vec: Improved mass spectral similarity scoring through learning of structural relationships]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765968472856-fb11732c-e83e-4b3d-aa54-8c4a964def85/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pcbi.1008724</link>
            <description><![CDATA[<p class="para" id="N65539">Spectral similarity is used as a proxy for structural similarity in many tandem mass spectrometry (MS/MS) based metabolomics analyses such as library matching and molecular networking. Although weaknesses in the relationship between spectral similarity scores and the true structural similarities have been described, little development of alternative scores has been undertaken. Here, we introduce Spec2Vec, a novel spectral similarity score inspired by a natural language processing algorithm—Word2Vec. Spec2Vec learns fragmental relationships within a large set of spectral data to derive abstract spectral embeddings that can be used to assess spectral similarities. Using data derived from GNPS MS/MS libraries including spectra for nearly 13,000 unique molecules, we show how Spec2Vec scores correlate better with structural similarity than cosine-based scores. We demonstrate the advantages of Spec2Vec in library matching and molecular networking. Spec2Vec is computationally more scalable allowing structural analogue searches in large databases within seconds.</p><p class="para" id="N65542">Most metabolomics analyses rely upon matching observed fragmentation mass spectra to library spectra for structural annotation or compare spectra with each other through network analysis. As a key part of such processes, scoring functions are used to assess the similarity between pairs of fragment spectra. No studies have so far proposed scores fundamentally different to the popular cosine-based similarity score, despite the fact that its limitations are well understood. We propose a novel spectral similarity score known as Spec2Vec which adapts algorithms from natural language processing to learn relationships between peaks from co-occurrences across large spectra datasets. We find that similarities computed with Spec2Vec i) correlate better to structural similarity than cosine-based scores, ii) subsequently gives better performance in library matching tasks, and iii) is computationally more scalable than cosine-based scores. Given the central place of similarity scoring in key metabolomics analysis tasks such as library matching and spectral networking, we expect Spec2Vec to make a broad impact in all fields that rely upon untargeted metabolomics.</p>]]></description>
            <pubDate><![CDATA[2021-02-16T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[A performance comparison of supervised machine learning models for Covid-19 tweets sentiment analysis]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765948169409-3a353d12-4ea9-46b9-80a5-e1768f89c08a/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245909</link>
            <description><![CDATA[<p class="para" id="N65539">The spread of Covid-19 has resulted in worldwide health concerns. Social media is increasingly used to share news and opinions about it. A realistic assessment of the situation is necessary to utilize resources optimally and appropriately. In this research, we perform Covid-19 tweets sentiment analysis using a supervised machine learning approach. Identification of Covid-19 sentiments from tweets would allow informed decisions for better handling the current pandemic situation. The used dataset is extracted from Twitter using IDs as provided by the IEEE data port. Tweets are extracted by an in-house built crawler that uses the Tweepy library. The dataset is cleaned using the preprocessing techniques and sentiments are extracted using the TextBlob library. The contribution of this work is the performance evaluation of various machine learning classifiers using our proposed feature set. This set is formed by concatenating the bag-of-words and the term frequency-inverse document frequency. Tweets are classified as positive, neutral, or negative. Performance of classifiers is evaluated on the accuracy, precision, recall, and <i>F</i><sub>1</sub> score. For completeness, further investigation is made on the dataset using the Long Short-Term Memory (LSTM) architecture of the deep learning model. The results show that Extra Trees Classifiers outperform all other models by achieving a 0.93 accuracy score using our proposed concatenated features set. The LSTM achieves low accuracy as compared to machine learning classifiers. To demonstrate the effectiveness of our proposed feature set, the results are compared with the Vader sentiment analysis technique based on the GloVe feature extraction approach.</p>]]></description>
            <pubDate><![CDATA[2021-02-25T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Unsupervised manifold learning of collective behavior]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765946511625-0ca05406-aefb-4aa7-b5a3-c002ffce52fd/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pcbi.1007811</link>
            <description><![CDATA[<p class="para" id="N65539">Collective behavior is an emergent property of numerous complex systems, from financial markets to cancer cells to predator-prey ecological systems. Characterizing modes of collective behavior is often done through human observation, training generative models, or other supervised learning techniques. Each of these cases requires knowledge of and a method for characterizing the macro-state(s) of the system. This presents a challenge for studying novel systems where there may be little prior knowledge. Here, we present a new unsupervised method of detecting emergent behavior in complex systems, and discerning between distinct collective behaviors. We require only metrics, <i>d</i><sup>(1)</sup>, <i>d</i><sup>(2)</sup>, defined on the set of agents, <i>X</i>, which measure agents’ nearness in variables of interest. We apply the method of diffusion maps to the systems (<i>X</i>, <i>d</i><sup>(<i>i</i>)</sup>) to recover efficient embeddings of their interaction networks. Comparing these geometries, we formulate a measure of similarity between two networks, called the map alignment statistic (MAS). A large MAS is evidence that the two networks are codetermined in some fashion, indicating an emergent relationship between the metrics <i>d</i><sup>(1)</sup> and <i>d</i><sup>(2)</sup>. Additionally, the form of the macro-scale organization is encoded in the covariances among the two sets of diffusion map components. Using these covariances we discern between different modes of collective behavior in a data-driven, unsupervised manner. This method is demonstrated on a synthetic flocking model as well as empirical fish schooling data. We show that our state classification subdivides the known behaviors of the school in a meaningful manner, leading to a finer description of the system’s behavior.</p><p class="para" id="N65542">Many complex systems in society and nature exhibit collective behavior where individuals’ local interactions lead to system-wide organization. One challenge we face today is to identify and characterize these emergent behaviors, and here we have developed a new method for analyzing data from individuals, to detect when a given complex system is exhibiting system-wide organization. Importantly, our approach requires no prior knowledge of the fashion in which the collective behavior arises, or the macro-scale variables in which it manifests. We apply the new method to an agent-based model and empirical observations of fish schooling. While we have demonstrated the utility of our approach to biological systems, it can be applied widely to financial, medical, and technological systems for example.</p>]]></description>
            <pubDate><![CDATA[2021-02-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[A dataset of human and robot approach behaviors into small free-standing conversational groups]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765944183000-e30083be-f790-4a1b-b89f-35c3327411dd/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247364</link>
            <description><![CDATA[<p class="para" id="N65539">The analysis and simulation of the interactions that occur in group situations is important when humans and artificial agents, physical or virtual, must coordinate when inhabiting similar spaces or even collaborate, as in the case of human-robot teams. Artificial systems should adapt to the natural interfaces of humans rather than the other way around. Such systems should be sensitive to human behaviors, which are often social in nature, and account for human capabilities when planning their own behaviors. A limiting factor relates to our understanding of how humans behave with respect to each other and with artificial embodiments, such as robots. To this end, we present <i>CongreG8</i> (pronounced ‘con-gre-gate’), a novel dataset containing the full-body motions of free-standing conversational groups of three humans and a newcomer that approaches the groups with the intent of joining them. The aim has been to collect an accurate and detailed set of positioning, orienting and full-body behaviors when a newcomer approaches and joins a small group. The dataset contains trials from human and robot newcomers. Additionally, it includes questionnaires about the personality of participants (BFI-10), their perception of robots (Godspeed), and custom human/robot interaction questions. An overview and analysis of the dataset is also provided, which suggests that human groups are more likely to alter their configuration to accommodate a human newcomer than a robot newcomer. We conclude by providing three use cases that the dataset has already been applied to in the domains of behavior detection and generation in real and virtual environments.</p><p class="para" id="N65544">A sample of the CongreG8 dataset is available at https://zenodo.org/record/4537811.</p>]]></description>
            <pubDate><![CDATA[2021-02-25T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Visual light perceptions caused by medical linear accelerator: Findings of machine-learning algorithms in a prospective questionnaire-based case–control study]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765944030034-4d3ecf4a-cf5c-4622-918b-571febd108fc/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247597</link>
            <description><![CDATA[<p class="para" id="N65539">This study aimed to investigate the possible incidence of visual light perceptions (VLPs) during radiation therapy (RT). We analyzed whether VLPs could be affected by differences in the radiation energy, prescription doses, age, sex, or RT locations, and whether all VLPs were caused by radiation. From November 2016 to August 2018, a total of 101 patients who underwent head-and-neck or brain RT were screened. After receiving RT, questionnaires were completed, and the subjects were interviewed. Random forests (RF), a tree-based machine learning algorithm, and logistic regression (LR) analyses were compared by the area under the curve (AUC), and the algorithm that achieved the highest AUC was selected. The dataset sample was based on treatment with non-human units, and a total of 293 treatment fields from 78 patients were analyzed. VLPs were detected only in 122 of the 293 exposure portals (40.16%). The dataset was randomly divided into 80% and 20% as the training set and test set, respectively. In the test set, RF achieved an AUC of 0.888, whereas LR achieved an AUC of 0.773. In this study, the retina fraction dose was the most important continuous variable and had a positive effect on VLP. Age was the most important categorical variable. In conclusion, the visual light perception phenomenon by the human body during RT is induced by radiation rather than being a self-suggested hallucination or induced by phosphenes.</p>]]></description>
            <pubDate><![CDATA[2021-02-25T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Monitoring social distancing under various low light conditions with deep learning and a single motionless time of flight camera]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765943238656-2b64b031-a0eb-4b0e-a58f-5b0d92fe5edd/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247440</link>
            <description><![CDATA[<p class="para" id="N65539">The purpose of this work is to provide an effective social distance monitoring solution in low light environments in a pandemic situation. The raging coronavirus disease 2019 (COVID-19) caused by the SARS-CoV-2 virus has brought a global crisis with its deadly spread all over the world. In the absence of an effective treatment and vaccine the efforts to control this pandemic strictly rely on personal preventive actions, e.g., handwashing, face mask usage, environmental cleaning, and most importantly on social distancing which is the only expedient approach to cope with this situation. Low light environments can become a problem in the spread of disease because of people’s night gatherings. Especially, in summers when the global temperature is at its peak, the situation can become more critical. Mostly, in cities where people have congested homes and no proper air cross-system is available. So, they find ways to get out of their homes with their families during the night to take fresh air. In such a situation, it is necessary to take effective measures to monitor the safety distance criteria to avoid more positive cases and to control the death toll. In this paper, a deep learning-based solution is proposed for the above-stated problem. The proposed framework utilizes the you only look once v4 (YOLO v4) model for real-time object detection and the social distance measuring approach is introduced with a single motionless time of flight (ToF) camera. The risk factor is indicated based on the calculated distance and safety distance violations are highlighted. Experimental results show that the proposed model exhibits good performance with 97.84% mean average precision (mAP) score and the observed mean absolute error (MAE) between actual and measured social distance values is 1.01 cm.</p>]]></description>
            <pubDate><![CDATA[2021-02-25T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[From heterogeneous healthcare data to disease-specific biomarker networks: A hierarchical Bayesian network approach]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765943199461-79a67ec5-e757-4111-b4ef-9c93f872d6a7/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pcbi.1008735</link>
            <description><![CDATA[<p class="para" id="N65539">In this work, we introduce an entirely data-driven and automated approach to reveal disease-associated biomarker and risk factor networks from heterogeneous and high-dimensional healthcare data. Our workflow is based on Bayesian networks, which are a popular tool for analyzing the interplay of biomarkers. Usually, data require extensive manual preprocessing and dimension reduction to allow for effective learning of Bayesian networks. For heterogeneous data, this preprocessing is hard to automatize and typically requires domain-specific prior knowledge. We here combine Bayesian network learning with hierarchical variable clustering in order to detect groups of similar features and learn interactions between them entirely automated. We present an optimization algorithm for the adaptive refinement of such group Bayesian networks to account for a specific target variable, like a disease. The combination of Bayesian networks, clustering, and refinement yields low-dimensional but disease-specific interaction networks. These networks provide easily interpretable, yet accurate models of biomarker interdependencies. We test our method extensively on simulated data, as well as on data from the Study of Health in Pomerania (SHIP-TREND), and demonstrate its effectiveness using non-alcoholic fatty liver disease and hypertension as examples. We show that the group network models outperform available biomarker scores, while at the same time, they provide an easily interpretable interaction network.</p><p class="para" id="N65542">High-dimensional and heterogeneous healthcare data, such as electronic health records or epidemiological study data, contain much information on yet unknown risk factors that are associated with disease development. The identification of these risk factors may help to improve prevention, diagnosis, and therapy. Bayesian networks are powerful statistical models that can decipher these complex relationships. However, high dimensionality and heterogeneity of data, together with missing values and high feature correlation, make it difficult to automatically learn a good model from data. To facilitate the use of network models, we present a novel, fully automated workflow that combines network learning with hierarchical clustering. The algorithm reveals groups of strongly related features and models the interactions among those groups. It results in simpler network models that are easier to analyze. We introduce a method of adaptive refinement of such models to ensure that disease-relevant parts of the network are modeled in great detail. Our approach makes it easy to learn compact, accurate, and easily interpretable biomarker interaction networks. We test our method extensively on simulated data as well as data from the Study of Health in Pomerania (SHIP-Trend) by learning models of hypertension and non-alcoholic fatty liver disease.</p>]]></description>
            <pubDate><![CDATA[2021-02-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Impact of between-tissue differences on pan-cancer predictions of drug sensitivity]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765942737539-78a59d12-8057-44f7-88bc-e4fe9ac64311/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pcbi.1008720</link>
            <description><![CDATA[<p class="para" id="N65539">Increased availability of drug response and genomics data for many tumor cell lines has accelerated the development of pan-cancer prediction models of drug response. However, it is unclear how much between-tissue differences in drug response and molecular characteristics may contribute to pan-cancer predictions. Also unknown is whether the performance of pan-cancer models could vary by cancer type. Here, we built a series of pan-cancer models using two datasets containing 346 and 504 cell lines, each with MEK inhibitor (MEKi) response and mRNA expression, point mutation, and copy number variation data, and found that, while the tissue-level drug responses are accurately predicted (between-tissue ρ = 0.88–0.98), only 5 of 10 cancer types showed successful within-tissue prediction performance (within-tissue ρ = 0.11–0.64). Between-tissue differences make substantial contributions to the performance of pan-cancer MEKi response predictions, as exclusion of between-tissue signals leads to a decrease in Spearman’s ρ from a range of 0.43–0.62 to 0.30–0.51. In practice, joint analysis of multiple cancer types usually has a larger sample size, hence greater power, than for one cancer type; and we observe that higher accuracy of pan-cancer prediction of MEKi response is almost entirely due to the sample size advantage. Success of pan-cancer prediction reveals how drug response in different cancers may invoke shared regulatory mechanisms despite tissue-specific routes of oncogenesis, yet predictions in different cancer types require flexible incorporation of between-cancer and within-cancer signals. As most datasets in genome sciences contain multiple levels of heterogeneity, careful parsing of group characteristics and within-group, individual variation is essential when making robust inference.</p><p class="para" id="N65542">One of the central goals for precision oncology is to tailor treatment of individual tumors by their molecular characteristics. While drug response predictions have traditionally been sought within each cancer type, it has long been hoped to develop more robust predictions by jointly considering diverse cancer types. While such pan-cancer approaches have improved in recent years, it remains unclear whether between-tissue differences are contributing to the reported pan-cancer prediction performance. This concern stems from the observation that, when cancer types differ in both molecular features and drug response, strong predictive information can come mainly from differences among tissue types. Our study finds that both between- and within-cancer type signals provide substantial contributions to pan-cancer drug response prediction models, and about half of the cancer types examined are poorly predicted despite strong overall performance across all cancer types. We also find that pan-cancer prediction models perform similarly or better than cancer type-specific models, and in many cases the advantage of pan-cancer models is due to the larger number of samples available for pan-cancer analysis. Our results highlight tissue-of-origin as a key consideration for pan-cancer drug response prediction models, and recommend cancer type-specific considerations when translating pan-cancer prediction models for clinical use.</p>]]></description>
            <pubDate><![CDATA[2021-02-25T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Revealing posturographic profile of patients with Parkinsonian syndromes through a novel hypothesis testing framework based on machine learning]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765942055090-97b0c0e6-4582-4c5e-8d35-0a2d9a922375/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246790</link>
            <description><![CDATA[<p class="para" id="N65539">Falling in Parkinsonian syndromes (PS) is associated with postural instability and consists a common cause of disability among PS patients. Current posturographic practices record the body’s center-of-pressure displacement (statokinesigram) while the patient stands on a force platform. Statokinesigrams, after appropriate processing, can offer numerous posturographic features. This fact, although beneficial, challenges the efforts for valid statistics via standard univariate approaches. In this work, 123 PS patients were classified into fallers (PS<sub>F</sub>) or non-faller (PS<sub>NF</sub>) based on the clinical assessment, and underwent simple Romberg Test (eyes open/eyes closed). We developed a non-parametric multivariate two-sample test (ts-AUC) based on machine learning, in order to examine statokinesigrams’ differences between PS<sub>F</sub> and PS<sub>NF</sub>. We analyzed posturographic features using both multiple testing with <i>p</i>-value adjustment and ts-AUC. While ts-AUC showed significant difference between groups (<i>p</i>-value = 0.01), multiple testing did not agree with this result (eyes open). PS<sub>F</sub> showed significantly increased antero-posterior movements as well as increased posturographic area compared to PS<sub>NF</sub>. Our study highlights the superiority of ts-AUC compared to standard statistical tools in distinguishing PS<sub>F</sub> and PS<sub>NF</sub> in multidimensional space. Machine learning-based statistical tests can be seen as a natural extension of classical statistics and should be considered, especially when dealing with multifactorial assessments.</p>]]></description>
            <pubDate><![CDATA[2021-02-25T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[TranSynergy: Mechanism-driven interpretable deep neural network for the synergistic prediction and pathway deconvolution of drug combinations]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765941989472-6e3cec9a-6e07-4439-9bf7-48a7f1dcc70d/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pcbi.1008653</link>
            <description><![CDATA[<p class="para" id="N65539">Drug combinations have demonstrated great potential in cancer treatments. They alleviate drug resistance and improve therapeutic efficacy. The fast-growing number of anti-cancer drugs has caused the experimental investigation of all drug combinations to become costly and time-consuming. Computational techniques can improve the efficiency of drug combination screening. Despite recent advances in applying machine learning to synergistic drug combination prediction, several challenges remain. First, the performance of existing methods is suboptimal. There is still much space for improvement. Second, biological knowledge has not been fully incorporated into the model. Finally, many models are lack interpretability, limiting their clinical applications. To address these challenges, we have developed a knowledge-enabled and self-attention transformer boosted deep learning model, TranSynergy, which improves the performance and interpretability of synergistic drug combination prediction. TranSynergy is designed so that the cellular effect of drug actions can be explicitly modeled through cell-line gene dependency, gene-gene interaction, and genome-wide drug-target interaction. A novel Shapley Additive Gene Set Enrichment Analysis (SA-GSEA) method has been developed to deconvolute genes that contribute to the synergistic drug combination and improve model interpretability. Extensive benchmark studies demonstrate that TranSynergy outperforms the state-of-the-art method, suggesting the potential of mechanism-driven machine learning. Novel pathways that are associated with the synergistic combinations are revealed and supported by experimental evidences. They may provide new insights into identifying biomarkers for precision medicine and discovering new anti-cancer therapies. Several new synergistic drug combinations have been predicted with high confidence for ovarian cancer which has few treatment options. The code is available at https://github.com/qiaoliuhub/drug_combination.</p><p class="para" id="N65542">The number of anti-cancer drugs has been consistently and quickly growing. They are mainly used as standardized mono-therapy. Drug combinations show substantial advantages over the anti-cancer mono-therapy. Cancer cells treated with the mono-therapy could later activate bypassing pathways and harbor drug resistances. Drug combinations can alleviate this issue by using a smaller doses of each anti-cancer drug or targeting multiple oncogenic pathways. However, the investigation of all anti-cancer drug combinations using experimental methods is costly and time-consuming. Machine learning provides an attractive solution to screening synergistic drug combinations, but it is a black-box and not easy to explain. We have developed a knowledge-enabled deep learning model, TranSynergy, to predict synergistic drug combinations and have demonstrated that our model outperformed other state-of-the-art methods. A novel Shapley Additive Gene Set Enrichment Analysis (SA-GSEA) method is introduced to improve the interpretability of the machine learning model. Using TransSynergy and SA-GSEA, we can deconvolute genes responsible for the synergistic drug combination, suggesting the potential of machine learning in developing precision anti-cancer therapy.</p>]]></description>
            <pubDate><![CDATA[2021-02-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[SOM-LWL method for identification of COVID-19 on chest X-rays]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765935812552-6d4b4bf3-7ea9-4419-ba13-c870e8f444ee/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247176</link>
            <description><![CDATA[<p class="para" id="N65539">The outbreak of coronavirus disease 2019 (COVID-19) has had an immense impact on world health and daily life in many countries. Sturdy observing of the initial site of infection in patients is crucial to gain control in the struggle with COVID-19. The early automated detection of the recent coronavirus disease (COVID-19) will help to limit its dissemination worldwide. Many initial studies have focused on the identification of the genetic material of coronavirus and have a poor detection rate for long-term surgery. The first imaging procedure that played an important role in COVID-19 treatment was the chest X-ray. Radiological imaging is often used as a method that emphasizes the performance of chest X-rays. Recent findings indicate the presence of COVID-19 in patients with irregular findings on chest X-rays. There are many reports on this topic that include machine learning strategies for the identification of COVID-19 using chest X-rays. Other current studies have used non-public datasets and complex artificial intelligence (AI) systems. In our research, we suggested a new COVID-19 identification technique based on the locality-weighted learning and self-organization map (LWL-SOM) strategy for detecting and capturing COVID-19 cases. We first grouped images from chest X-ray datasets based on their similar features in different clusters using the SOM strategy in order to discriminate between the COVID-19 and non-COVID-19 cases. Then, we built our intelligent learning model based on the LWL algorithm to diagnose and detect COVID-19 cases. The proposed SOM-LWL model improved the correlation coefficient performance results between the Covid19, no-finding, and pneumonia cases; pneumonia and no-finding cases; Covid19 and pneumonia cases; and Covid19 and no-finding cases from 0.9613 to 0.9788, 0.6113 to 1 0.8783 to 0.9999, and 0.8894 to 1, respectively. The proposed LWL-SOM had better results for discriminating COVID-19 and non-COVID-19 patients than the current machine learning-based solutions using AI evaluation measures.</p>]]></description>
            <pubDate><![CDATA[2021-02-24T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Blood biomarker discovery for autism spectrum disorder: A proteomic analysis]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765935714832-03c0676e-e1fb-45b4-9d5b-859f8411ec3e/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246581</link>
            <description><![CDATA[<p class="para" id="N65539">Autism spectrum disorder (ASD) is a neurodevelopmental disorder characterized by deficits in social communication and social interaction and restricted, repetitive patterns of behavior, interests, or activities. Given the lack of specific pharmacological therapy for ASD and the clinical heterogeneity of the disorder, current biomarker research efforts are geared mainly toward identifying markers for determining ASD risk or for assisting with a diagnosis. A wide range of putative biological markers for ASD is currently being investigated. Proteomic analyses indicate that the levels of many proteins in plasma/serum are altered in ASD, suggesting that a panel of proteins may provide a blood biomarker for ASD. Serum samples from 76 boys with ASD and 78 typically developing (TD) boys, 18 months-8 years of age, were analyzed to identify possible early biological markers for ASD. Proteomic analysis of serum was performed using SomaLogic’s SOMAScan<sup>TM</sup> assay 1.3K platform. A total of 1,125 proteins were analyzed. There were 86 downregulated proteins and 52 upregulated proteins in ASD (FDR &lt; 0.05). Combining three different algorithms, we found a panel of 9 proteins that identified ASD with an area under the curve (AUC) = 0.8599±0.0640, with specificity and sensitivity of 0.8217±0.1178 and 0.835±0.1176, respectively. All 9 proteins were significantly different in ASD compared with TD boys, and were significantly correlated with ASD severity as measured by ADOS total scores. Using machine learning methods, a panel of serum proteins was identified that may be useful as a blood biomarker for ASD in boys. Further verification of the protein biomarker panel with independent test sets is warranted.</p>]]></description>
            <pubDate><![CDATA[2021-02-24T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Mental fatigue prediction during eye-typing]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765935570638-d2fad3a8-e9d2-4e65-93f8-fb3a875ff649/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246739</link>
            <description><![CDATA[<p class="para" id="N65539">Mental fatigue is a common problem associated with neurological disorders. Until now, there has not been a method to assess mental fatigue on a continuous scale. Camera-based eye-typing is commonly used for communication by people with severe neurological disorders. We designed a working memory-based eye-typing experiment with 18 healthy participants, and obtained eye-tracking and typing performance data in addition to their subjective scores on perceived effort for every sentence typed and mental fatigue, to create a model of mental fatigue for eye-typing. The features of the model were the eye-based blink frequency, eye height and baseline-related pupil diameter. We predicted subjective ratings of mental fatigue on a six-point Likert scale, using random forest regression, with 22% lower mean absolute error than using simulations. When additionally including task difficulty (i.e. the difficulty of the sentences typed) as a feature, the variance explained by the model increased by 9%. This indicates that task difficulty plays an important role in modelling mental fatigue. The results demonstrate the feasibility of objective and non-intrusive measurement of fatigue on a continuous scale.</p>]]></description>
            <pubDate><![CDATA[2021-02-22T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Identification of RNA pseudouridine sites using deep learning approaches]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765934799121-dea228b0-5a78-402a-b21f-a19a6ccaf4ea/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247511</link>
            <description><![CDATA[<p class="para" id="N65539">Pseudouridine(Ψ) is widely popular among various RNA modifications which have been confirmed to occur in rRNA, mRNA, tRNA, and nuclear/nucleolar RNA. Hence, identifying them has vital significance in academic research, drug development and gene therapies. Several laboratory techniques for Ψ identification have been introduced over the years. Although these techniques produce satisfactory results, they are costly, time-consuming and requires skilled experience. As the lengths of RNA sequences are getting longer day by day, an efficient method for identifying pseudouridine sites using computational approaches is very important. In this paper, we proposed a multi-channel convolution neural network using binary encoding. We employed k-fold cross-validation and grid search to tune the hyperparameters. We evaluated its performance in the independent datasets and found promising results. The results proved that our method can be used to identify pseudouridine sites for associated purposes. We have also implemented an easily accessible web server at http://103.99.176.239/ipseumulticnn/.</p>]]></description>
            <pubDate><![CDATA[2021-02-23T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Review of machine learning methods in soft robotics]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765925299050-11ad9bd6-1aac-4dd8-b9bf-6aa70b3793b9/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246102</link>
            <description><![CDATA[<p class="para" id="N65539">Soft robots have been extensively researched due to their flexible, deformable, and adaptive characteristics. However, compared to rigid robots, soft robots have issues in modeling, calibration, and control in that the innate characteristics of the soft materials can cause complex behaviors due to non-linearity and hysteresis. To overcome these limitations, recent studies have applied various approaches based on machine learning. This paper presents existing machine learning techniques in the soft robotic fields and categorizes the implementation of machine learning approaches in different soft robotic applications, which include soft sensors, soft actuators, and applications such as soft wearable robots. An analysis of the trends of different machine learning approaches with respect to different types of soft robot applications is presented; in addition to the current limitations in the research field, followed by a summary of the existing machine learning methods for soft robots.</p>]]></description>
            <pubDate><![CDATA[2021-02-18T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[An artificial neural network approach to detect presence and severity of Parkinson’s disease via gait parameters]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765924884103-927abb24-b32a-4dc3-a6e6-2a7b97867dae/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244396</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Introduction</h3><p class="para" id="N65543">Gait deficits are debilitating in people with Parkinson’s disease (PwPD), which inevitably deteriorate over time. Gait analysis is a valuable method to assess disease-specific gait patterns and their relationship with the clinical features and progression of the disease.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Objectives</h3><p class="para" id="N65549">Our study aimed to i) develop an automated diagnostic algorithm based on machine-learning techniques (artificial neural networks [ANNs]) to classify the gait deficits of PwPD according to disease progression in the Hoehn and Yahr (H-Y) staging system, and ii) identify a minimum set of gait classifiers.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Methods</h3><p class="para" id="N65555">We evaluated 76 PwPD (H-Y stage 1–4) and 67 healthy controls (HCs) by computerized gait analysis. We computed the time-distance parameters and the ranges of angular motion (RoMs) of the hip, knee, ankle, trunk, and pelvis. Principal component analysis was used to define a subset of features including all gait variables. An ANN approach was used to identify gait deficits according to the H-Y stage.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Results</h3><p class="para" id="N65561">We identified a combination of a small number of features that distinguished PwPDs from HCs (one combination of two features: knee and trunk rotation RoMs) and identified the gait patterns between different H-Y stages (two combinations of four features: walking speed and hip, knee, and ankle RoMs; walking speed and hip, knee, and trunk rotation RoMs).</p></div><div class="section" id="sec005"><h3 class="BHead" id="nov000-5">Conclusion</h3><p class="para" id="N65567">The ANN approach enabled automated diagnosis of gait deficits in several symptomatic stages of Parkinson’s disease. These results will inspire future studies to test the utility of gait classifiers for the evaluation of treatments that could modify disease progression.</p></div>]]></description>
            <pubDate><![CDATA[2021-02-19T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Dynamic graph embedding for outlier detection on multiple meteorological time series]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765924662624-4a70aa12-57bc-42ff-b310-6ed87ae4eecb/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247119</link>
            <description><![CDATA[<p class="para" id="N65539">Existing dynamic graph embedding-based outlier detection methods mainly focus on the evolution of graphs and ignore the similarities among them. To overcome this limitation for the effective detection of abnormal climatic events from meteorological time series, we proposed a dynamic graph embedding model based on graph proximity, called DynGPE. Climatic events are represented as a graph where each vertex indicates meteorological data and each edge indicates a spurious relationship between two meteorological time series that are not causally related. The graph proximity is described as the distance between two graphs. DynGPE can cluster similar climatic events in the embedding space. Abnormal climatic events are distant from most of the other events and can be detected using outlier detection methods. We conducted experiments by applying three outlier detection methods (i.e., isolation forest, local outlier factor, and box plot) to real meteorological data. The results showed that DynGPE achieves better results than the baseline by 44.3% on average in terms of the F-measure. Isolation forest provides the best performance and stability. It achieved higher results than the local outlier factor and box plot methods, namely, by 15.4% and 78.9% on average, respectively.</p>]]></description>
            <pubDate><![CDATA[2021-02-18T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Context-sensitive smart glasses monitoring wear position and activity for therapy compliance—A proof of concept]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765923563773-284170de-c468-47c5-9163-ddab6cf47cbe/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247389</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Purpose</h3><p class="para" id="N65543">To improve the acceptance and compliance of treatment of amblyopia, the aim of this study was to show that it is feasible to design an electronic frame for context-sensitive liquid crystal glasses, which can measure the state of wear position in a robust manner and detect distinct motion patterns for activity recognition.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Methods</h3><p class="para" id="N65549">Different temple designs with integrated temperature and capacitive sensors were developed to realize the detection of the state of wear position to distinguish three states (correct position/wrong position/glasses taken off). The electronic glasses frame was further designed as a tool for accelerometer data acquisition, which was used for algorithm development for activity classification. For this purpose, training data of 20 voluntary healthy adult subjects (5 females, 15 males) were recorded and a 10-fold cross-validation was computed for classifier selection. In order to perform functional testing of the electronic glasses frame, a proof of concept study was performed in a small group of healthy adults. Four healthy adult subjects (2 females, 2 males) were included to wear the electronic glasses frame and to protocol their activities in their everyday life according to a defined test protocol. Individual and averaged results for the precision of the state of wear position detection and of the activity recognition were calculated.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Results</h3><p class="para" id="N65555">Context-sensitive control algorithms were developed which detected the state of wear position and activity in a proof of concept. The pilot study revealed an average of 91.4% agreement of the detected states of wear position. The activity recognition match was 82.2% when applying an additional filter criterion. Removing the glasses was always detected 100% correctly.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Conclusion</h3><p class="para" id="N65561">The principles investigated are suitable for detecting the glasses’ state of wear position and for recognizing the wearer´s activity in a smart glasses concept.</p></div>]]></description>
            <pubDate><![CDATA[2021-02-19T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[A machine learning approach to identify distinct subgroups of veterans at risk for hospitalization or death using administrative and electronic health record data]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765923505319-b9e7b5f4-5292-4b3d-b89a-5e25dcc82e0d/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247203</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Background</h3><p class="para" id="N65543">Identifying individuals at risk for future hospitalization or death has been a major priority of population health management strategies. High-risk individuals are a heterogeneous group, and existing studies describing heterogeneity in high-risk individuals have been limited by data focused on clinical comorbidities and not socioeconomic or behavioral factors. We used machine learning clustering methods and linked comorbidity-based, sociodemographic, and psychobehavioral data to identify subgroups of high-risk Veterans and study long-term outcomes, hypothesizing that factors other than comorbidities would characterize several subgroups.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Methods and findings</h3><p class="para" id="N65549">In this cross-sectional study, we used data from the VA Corporate Data Warehouse, a national repository of VA administrative claims and electronic health data. To identify high-risk Veterans, we used the Care Assessment Needs (CAN) score, a routinely-used VA model that predicts a patient’s percentile risk of hospitalization or death at one year. Our study population consisted of 110,000 Veterans who were randomly sampled from 1,920,436 Veterans with a CAN score≥75<sup>th</sup> percentile in 2014. We categorized patient-level data into 119 independent variables based on demographics, comorbidities, pharmacy, vital signs, laboratories, and prior utilization. We used a previously validated density-based clustering algorithm to identify 30 subgroups of high-risk Veterans ranging in size from 50 to 2,446 patients. Mean CAN score ranged from 72.4 to 90.3 among subgroups. Two-year mortality ranged from 0.9% to 45.6% and was highest in the home-based care and metastatic cancer subgroups. Mean inpatient days ranged from 1.4 to 30.5 and were highest in the post-surgery and blood loss anemia subgroups. Mean emergency room visits ranged from 1.0 to 4.3 and were highest in the chronic sedative use and polysubstance use with amphetamine predominance subgroups. Five subgroups were distinguished by psychobehavioral factors and four subgroups were distinguished by sociodemographic factors.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Conclusions</h3><p class="para" id="N65558">High-risk Veterans are a heterogeneous population consisting of multiple distinct subgroups–many of which are not defined by clinical comorbidities–with distinct utilization and outcome patterns. To our knowledge, this represents the largest application of ML clustering methods to subgroup a high-risk population. Further study is needed to determine whether distinct subgroups may benefit from individualized interventions.</p></div>]]></description>
            <pubDate><![CDATA[2021-02-19T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Secondary bile acid ursodeoxycholic acid alters weight, the gut microbiota, and the bile acid pool in conventional mice]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765923482261-00888a19-80a7-4222-9934-873d3ce8c7c4/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246161</link>
            <description><![CDATA[<p class="para" id="N65539">Ursodeoxycholic acid (commercially available as ursodiol) is a naturally occurring bile acid that is used to treat a variety of hepatic and gastrointestinal diseases. Ursodiol can modulate bile acid pools, which have the potential to alter the gut microbiota community structure. In turn, the gut microbial community can modulate bile acid pools, thus highlighting the interconnectedness of the gut microbiota-bile acid-host axis. Despite these interactions, it remains unclear if and how exogenously administered ursodiol shapes the gut microbial community structure and bile acid pool in conventional mice. This study aims to characterize how ursodiol alters the gastrointestinal ecosystem in conventional mice. C57BL/6J wildtype mice were given one of three doses of ursodiol (50, 150, or 450 mg/kg/day) by oral gavage for 21 days. Alterations in the gut microbiota and bile acids were examined including stool, ileal, and cecal content. Bile acids were also measured in serum. Significant weight loss was seen in mice treated with the low and high dose of ursodiol. Alterations in the microbial community structure and bile acid pool were seen in ileal and cecal content compared to pretreatment, and longitudinally in feces following the 21-day ursodiol treatment. In both ileal and cecal content, members of the Lachnospiraceae Family significantly contributed to the changes observed. This study is the first to provide a comprehensive view of how exogenously administered ursodiol shapes the healthy gastrointestinal ecosystem in conventional mice. Further studies to investigate how these changes in turn modify the host physiologic response are important.</p>]]></description>
            <pubDate><![CDATA[2021-02-18T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[DTI-SNNFRA: Drug-target interaction prediction by shared nearest neighbors and fuzzy-rough approximation]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765922577716-3b5c8a07-0e25-4418-9950-1ba1449a0def/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246920</link>
            <description><![CDATA[<p class="para" id="N65539"><i>In-silico</i> prediction of repurposable drugs is an effective drug discovery strategy that supplements <i>de-nevo</i> drug discovery from scratch. Reduced development time, less cost and absence of severe side effects are significant advantages of using drug repositioning. Most recent and most advanced artificial intelligence (AI) approaches have boosted drug repurposing in terms of throughput and accuracy enormously. However, with the growing number of drugs, targets and their massive interactions produce imbalanced data which may not be suitable as input to the classification model directly. Here, we have proposed DTI-SNNFRA, a framework for predicting drug-target interaction (DTI), based on shared nearest neighbour (SNN) and fuzzy-rough approximation (FRA). It uses sampling techniques to collectively reduce the vast search space covering the available drugs, targets and millions of interactions between them. DTI-SNNFRA operates in two stages: first, it uses SNN followed by a partitioning clustering for sampling the search space. Next, it computes the degree of fuzzy-rough approximations and proper degree threshold selection for the negative samples’ undersampling from all possible interaction pairs between drugs and targets obtained in the first stage. Finally, classification is performed using the positive and selected negative samples. We have evaluated the efficacy of DTI-SNNFRA using AUC (Area under ROC Curve), Geometric Mean, and F1 Score. The model performs exceptionally well with a high prediction score of 0.95 for ROC-AUC. The predicted drug-target interactions are validated through an existing drug-target database (Connectivity Map (Cmap)).</p>]]></description>
            <pubDate><![CDATA[2021-02-19T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Facial geometric feature extraction based emotional expression classification using machine learning algorithms]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765922247351-59009b91-57e9-48d7-93d6-48619d614041/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247131</link>
            <description><![CDATA[<p class="para" id="N65539">Emotion plays a significant role in interpersonal communication and also improving social life. In recent years, facial emotion recognition is highly adopted in developing human-computer interfaces (HCI) and humanoid robots. In this work, a triangulation method for extracting a novel set of geometric features is proposed to classify six emotional expressions (sadness, anger, fear, surprise, disgust, and happiness) using computer-generated markers. The subject’s face is recognized by using Haar-like features. A mathematical model has been applied to positions of eight virtual markers in a defined location on the subject’s face in an automated way. Five triangles are formed by manipulating eight markers’ positions as an edge of each triangle. Later, these eight markers are uninterruptedly tracked by Lucas- Kanade optical flow algorithm while subjects’ articulating facial expressions. The movement of the markers during facial expression directly changes the property of each triangle. The area of the triangle (AoT), Inscribed circle circumference (ICC), and the Inscribed circle area of a triangle (ICAT) are extracted as features to classify the facial emotions. These features are used to distinguish six different facial emotions using various types of machine learning algorithms. The inscribed circle area of the triangle (ICAT) feature gives a maximum mean classification rate of 98.17% using a Random Forest (RF) classifier compared to other features and classifiers in distinguishing emotional expressions.</p>]]></description>
            <pubDate><![CDATA[2021-02-18T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Feasibility of integrating canine olfaction with chemical and microbial profiling of urine to detect lethal prostate cancer]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765907356418-0b8e0559-2fb0-444d-ad5b-44a2ec2c3ef2/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245530</link>
            <description><![CDATA[<p class="para" id="N65539">Prostate cancer is the second leading cause of cancer death in men in the developed world. A more sensitive and specific detection strategy for lethal prostate cancer beyond serum prostate specific antigen (PSA) population screening is urgently needed. Diagnosis by canine olfaction, using dogs trained to detect cancer by smell, has been shown to be both specific and sensitive. While dogs themselves are impractical as scalable diagnostic sensors, machine olfaction for cancer detection is testable. However, studies bridging the divide between clinical diagnostic techniques, artificial intelligence, and molecular analysis remains difficult due to the significant divide between these disciplines. We tested the clinical feasibility of a cross-disciplinary, integrative approach to early prostate cancer biosensing in urine using trained canine olfaction, volatile organic compound (VOC) analysis by gas chromatography-mass spectroscopy (GC-MS) artificial neural network (ANN)-assisted examination, and microbial profiling in a double-blinded pilot study. Two dogs were trained to detect Gleason 9 prostate cancer in urine collected from biopsy-confirmed patients. Biopsy-negative controls were used to assess canine specificity as prostate cancer biodetectors. Urine samples were simultaneously analyzed for their VOC content in headspace via GC-MS and urinary microbiota content via 16S rDNA Illumina sequencing. In addition, the dogs’ diagnoses were used to train an ANN to detect significant peaks in the GC-MS data. The canine olfaction system was 71% sensitive and between 70–76% specific at detecting Gleason 9 prostate cancer. We have also confirmed VOC differences by GC-MS and microbiota differences by 16S rDNA sequencing between cancer positive and biopsy-negative controls. Furthermore, the trained ANN identified regions of interest in the GC-MS data, informed by the canine diagnoses. Methodology and feasibility are established to inform larger-scale studies using canine olfaction, urinary VOCs, and urinary microbiota profiling to develop machine olfaction diagnostic tools. Scalable multi-disciplinary tools may then be compared to PSA screening for earlier, non-invasive, more specific and sensitive detection of clinically aggressive prostate cancers in urine samples.</p>]]></description>
            <pubDate><![CDATA[2021-02-17T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[High density optical neuroimaging predicts surgeons’s subjective experience and skill levels]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765906148860-ea302bbb-cf4b-4841-9521-ed6f40e1ba0f/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247117</link>
            <description><![CDATA[<p class="para" id="N65539">Measuring cognitive load is important for surgical education and patient safety. Traditional approaches of measuring cognitive load of surgeons utilise behavioural metrics to measure performance and surveys and questionnaires to collect reports of subjective experience. These have disadvantages such as sporadic data, occasionally intrusive methodologies, subjective or misleading self-reporting. In addition, traditional approaches use subjective metrics that cannot distinguish between skill levels. Functional neuroimaging data was collected using a high density, wireless NIRS device from sixteen surgeons (11 attending surgeons and 5 surgery resident) and 17 students while they performed two laparoscopic tasks (Peg transfer and String pass). Participant’s subjective mental load was assessed using the NASA-TLX survey. Machine learning approaches were used for predicting the subjective experience and skill levels. The Prefrontal cortex (PFC) activations were greater in students who reported higher-than-median task load, as measured by the NASA-TLX survey. However in the case of attending surgeons the opposite tendency was observed, namely higher activations in the lower v higher task loaded subjects. We found that response was greater in the left PFC of students particularly near the dorso- and ventrolateral areas. We quantified the ability of PFC activation to predict the differences in skill and task load using machine learning while focussing on the effects of NIRS channel separation distance on the results. Our results showed that the classification of skill level and subjective task load could be predicted based on PFC activation with an accuracy of nearly 90%. Our finding shows that there is sufficient information available in the optical signals to make accurate predictions about the surgeons’ subjective experiences and skill levels. The high accuracy of results is encouraging and suggest the integration of the strategy developed in this study as a promising approach to design automated, more accurate and objective evaluation methods.</p>]]></description>
            <pubDate><![CDATA[2021-02-18T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Comparison of diagnostic performance between convolutional neural networks and human endoscopists for diagnosis of colorectal polyp: A systematic review and meta-analysis]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765900881352-d7b04fd5-2d5b-4977-af9e-1e5ba6ad5a3e/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246892</link>
            <description><![CDATA[<p class="para" id="N65539">Prospective randomized trials and observational studies have revealed that early detection, classification, and removal of neoplastic colorectal polyp (CP) significantly improve the prevention of colorectal cancer (CRC). The current effectiveness of the diagnostic performance of colonoscopy remains unsatisfactory with unstable accuracy. The convolutional neural networks (CNN) system based on artificial intelligence (AI) technology has demonstrated its potential to help endoscopists in increasing diagnostic accuracy. Nonetheless, several limitations of the CNN system and controversies exist on whether it provides a better diagnostic performance compared to human endoscopists. Therefore, this study sought to address this issue. Online databases (PubMed, Web of Science, Cochrane Library, and EMBASE) were used to search for studies conducted up to April 2020. Besides, the quality assessment of diagnostic accuracy scale-2 (QUADAS-2) was used to evaluate the quality of the enrolled studies. Moreover, publication bias was determined using the Deeks’ funnel plot. In total, 13 studies were enrolled for this meta-analysis (ranged between 2016 and 2020). Consequently, the CNN system had a satisfactory diagnostic performance in the field of CP detection (sensitivity: 0.848 [95% CI: 0.692–0.932]; specificity: 0.965 [95% CI: 0.946–0.977]; and AUC: 0.98 [95% CI: 0.96–0.99]) and CP classification (sensitivity: 0.943 [95% CI: 0.927–0.955]; specificity: 0.894 [95% CI: 0.631–0.977]; and AUC: 0.95 [95% CI: 0.93–0.97]). In comparison with human endoscopists, the CNN system was comparable to the expert but significantly better than the non-expert in the field of CP classification (CNN <i>vs</i>. expert: RDOR: 1.03, <i>P</i> = 0.9654; non-expert <i>vs</i>. expert: RDOR: 0.29, <i>P</i> = 0.0559; non-expert <i>vs</i>. CNN: 0.18, <i>P</i> = 0.0342). Therefore, the CNN system exhibited a satisfactory diagnostic performance for CP and could be used as a potential clinical diagnostic tool during colonoscopy.</p>]]></description>
            <pubDate><![CDATA[2021-02-16T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Patient journey through cases of depression from claims database using machine learning algorithms]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765900313706-76cc6f14-80ff-4a62-bb8e-f3a6fd27dd06/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247059</link>
            <description><![CDATA[<p class="para" id="N65539">Health insurance and acute hospital-based claims have recently become available as real-world data after marketing in Japan and, thus, classification and prediction using the machine learning approach can be applied to them. However, the methodology used for the analysis of real-world data has been hitherto under debate and research on visualizing the patient journey is still inconclusive. So far, to classify diseases based on medical histories and patient demographic background and to predict the patient prognosis for each disease, the correlation structure of real-world data has been estimated by machine learning. Therefore, we applied association analysis to real-world data to consider a combination of disease events as the patient journey for depression diagnoses. However, association analysis makes it difficult to interpret multiple outcome measures simultaneously and comprehensively. To address this issue, we applied the Topological Data Analysis (TDA) Mapper to sequentially interpret multiple indices, thus obtaining a visual classification of the diseases commonly associated with depression. Under this approach, the visual and continuous classification of related diseases may contribute to precision medicine research and can help pharmaceutical companies provide appropriate personalized medical care.</p>]]></description>
            <pubDate><![CDATA[2021-02-16T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Modeling vehicle ownership with machine learning techniques in the Greater Tamale Area, Ghana]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765900219433-8387c411-0634-4b58-ad06-d74fb1c69b26/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246044</link>
            <description><![CDATA[<p class="para" id="N65539">Vehicle ownership modeling and prediction is a crucial task in the transportation planning processes which, traditionally, uses statistical models in the modeling process. However, with the advancement in computing power of computers and Artificial Intelligence, Machine Learning (ML) algorithms are becoming an alternative or a complement to the statistical models in modeling the transportation planning processes. Although the application of ML algorithms to the transportation planning processes—like mode choice, and traffic forecasting and demand modeling—have received much attention in research and abound in literature, scanty attention is paid to its application to vehicle ownership modeling especially in the context of small to medium cities in developing countries. Therefore, this study attempts to fill this gap by modeling vehicle ownership in the Greater Tamale Area (GTA), a typically small to medium city in Ghana. Using a cross sectional survey of formal sectors workers, data was collected between June–August 2018. The study applied nine different ML classification algorithms to the dataset using 10-fold cross-validation technique/s and the Cohen-Kappa static/statistic to evaluate the predictive performance of each of the algorithms, and the Permutation Feature Importance to examine the features that contribute significantly to the prediction of vehicle ownership in GTA. The results showed that Linear Support Vector Classification (LinearSVC) classifier performed well in comparison with the other classifiers with regards to the overall predictive ability of the classifiers. In terms of class predictions, K- Nearest Neighbors (KNN) classifier performs well for no-vehicle class whiles Linear Support Vector Classification (LinearSVC) and GaussianNB classifiers performs well for motorcycle ownership. LinearSVC and Logistic Regression classifiers performed well on the car ownership class. Also, the results indicated that travel mode choice, average monthly income, average travel distance to workplace, average monthly expenditure on transport, duration of travel to workplace, occupational rank, age, household size and marital status were significant in predicting vehicle ownership for most of the classifiers. These findings could help policies makers carve out strategies that would reduce vehicle ownership but improve personal mobility.</p>]]></description>
            <pubDate><![CDATA[2021-02-16T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Functional parcellation of mouse visual cortex using statistical techniques reveals response-dependent clustering of cortical processing areas]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765899871460-050e8b2e-769a-4688-8432-3f9e5f9801af/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pcbi.1008548</link>
            <description><![CDATA[<p class="para" id="N65539">The visual cortex of the mouse brain can be divided into ten or more areas that each contain complete or partial retinotopic maps of the contralateral visual field. It is generally assumed that these areas represent discrete processing regions. In contrast to the conventional input-output characterizations of neuronal responses to standard visual stimuli, here we asked whether six of the core visual areas have responses that are functionally distinct from each other for a given visual stimulus set, by applying machine learning techniques to distinguish the areas based on their activity patterns. Visual areas defined by retinotopic mapping were examined using supervised classifiers applied to responses elicited by a range of stimuli. Using two distinct datasets obtained using wide-field and two-photon imaging, we show that the area labels predicted by the classifiers were highly consistent with the labels obtained using retinotopy. Furthermore, the classifiers were able to model the boundaries of visual areas using resting state cortical responses obtained without any overt stimulus, in both datasets. With the wide-field dataset, clustering neuronal responses using a constrained semi-supervised classifier showed graceful degradation of accuracy. The results suggest that responses from visual cortical areas can be classified effectively using data-driven models. These responses likely reflect unique circuits within each area that give rise to activity with stronger intra-areal than inter-areal correlations, and their responses to controlled visual stimuli across trials drive higher areal classification accuracy than resting state responses.</p><p class="para" id="N65542">The visual cortex has a prominent role in the processing of visual information by the brain. Previous work has segmented the mouse visual cortex into different areas based on the organization of retinotopic maps. Here, we collect responses of the visual cortex to various types of stimuli and ask if we could discover unique clusters from this dataset using machine learning methods. The retinotopy based area borders are used as ground truth to compare the performance of our clustering algorithms. We show our results on two datasets, one collected by the authors using wide-field imaging and another a publicly available dataset collected using two-photon imaging. The proposed supervised approach is able to predict the area labels accurately using neuronal responses to various visual stimuli. Following up on these results using visual stimuli, we hypothesized that each area of the mouse brain has unique responses that can be used to classify the area independently of stimuli. Experiments using resting state responses, without any overt stimulus, confirm this hypothesis. Such activity-based segmentation of the mouse visual cortex suggests that large-scale imaging combined with a machine learning algorithm may enable new insights into the functional organization of the visual cortex in mice and other species.</p>]]></description>
            <pubDate><![CDATA[2021-02-04T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Target spike patterns enable efficient and biologically plausible learning for complex temporal tasks]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765899313728-12301777-e75c-4ef9-93af-0dada8942321/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0247014</link>
            <description><![CDATA[<p class="para" id="N65539">Recurrent spiking neural networks (RSNN) in the brain learn to perform a wide range of perceptual, cognitive and motor tasks very efficiently in terms of energy consumption and their training requires very few examples. This motivates the search for biologically inspired learning rules for RSNNs, aiming to improve our understanding of brain computation and the efficiency of artificial intelligence. Several spiking models and learning rules have been proposed, but it remains a challenge to design RSNNs whose learning relies on biologically plausible mechanisms and are capable of solving complex temporal tasks. In this paper, we derive a learning rule, local to the synapse, from a simple mathematical principle, the maximization of the likelihood for the network to solve a specific task. We propose a novel target-based learning scheme in which the learning rule derived from likelihood maximization is used to mimic a specific spatio-temporal spike pattern that encodes the solution to complex temporal tasks. This method makes the learning extremely rapid and precise, outperforming state of the art algorithms for RSNNs. While error-based approaches, (e.g. e-prop) trial after trial optimize the internal sequence of spikes in order to progressively minimize the MSE we assume that a signal randomly projected from an external origin (e.g. from other brain areas) directly defines the target sequence. This facilitates the learning procedure since the network is trained from the beginning to reproduce the desired internal sequence. We propose two versions of our learning rule: spike-dependent and voltage-dependent. We find that the latter provides remarkable benefits in terms of learning speed and robustness to noise. We demonstrate the capacity of our model to tackle several problems like learning multidimensional trajectories and solving the classical temporal XOR benchmark. Finally, we show that an online approximation of the gradient ascent, in addition to guaranteeing complete locality in time and space, allows learning after very few presentations of the target output. Our model can be applied to different types of biological neurons. The analytically derived plasticity learning rule is specific to each neuron model and can produce a theoretical prediction for experimental validation.</p>]]></description>
            <pubDate><![CDATA[2021-02-16T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[HDSI: High dimensional selection with interactions algorithm on feature selection and testing]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765899290501-1a8ffc22-7b3e-4099-8d3e-7ad532f5ddab/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246159</link>
            <description><![CDATA[<p class="para" id="N65539">Feature selection on high dimensional data along with the interaction effects is a critical challenge for classical statistical learning techniques. Existing feature selection algorithms such as random LASSO leverages LASSO capability to handle high dimensional data. However, the technique has two main limitations, namely the inability to consider interaction terms and the lack of a statistical test for determining the significance of selected features. This study proposes a High Dimensional Selection with Interactions (HDSI) algorithm, a new feature selection method, which can handle high-dimensional data, incorporate interaction terms, provide the statistical inferences of selected features and leverage the capability of existing classical statistical techniques. The method allows the application of any statistical technique like LASSO and subset selection on multiple bootstrapped samples; each contains randomly selected features. Each bootstrap data incorporates interaction terms for the randomly sampled features. The selected features from each model are pooled and their statistical significance is determined. The selected statistically significant features are used as the final output of the approach, whose final coefficients are estimated using appropriate statistical techniques. The performance of HDSI is evaluated using both simulated data and real studies. In general, HDSI outperforms the commonly used algorithms such as LASSO, subset selection, adaptive LASSO, random LASSO and group LASSO.</p>]]></description>
            <pubDate><![CDATA[2021-02-16T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[PI Prob: A risk prediction and clinical guidance system for evaluating patients with recurrent infections]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765899065137-77aa59d0-3d50-479e-be88-a0240cda8305/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0237285</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Background</h3><p class="para" id="N65543">Primary immunodeficiency diseases represent an expanding set of heterogeneous conditions which are difficult to recognize clinically. Diagnostic rates outside of the newborn period have not changed appreciably. This concern underscores a need for novel methods of disease detection.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Objective</h3><p class="para" id="N65549">We built a Bayesian network to provide real-time risk assessment about primary immunodeficiency and to facilitate prescriptive analytics for initiating the most appropriate diagnostic work up. Our goal is to improve diagnostic rates for primary immunodeficiency and shorten time to diagnosis. We aimed to use readily available health record data and a small training dataset to prove utility in diagnosing patients with relatively rare features.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Methods</h3><p class="para" id="N65555">We extracted data from the Texas Children’s Hospital electronic health record on a large population of primary immunodeficiency patients (n = 1762) and appropriately-matched set of controls (n = 1698). From the cohorts, clinically relevant prior probabilities were calculated enabling construction of a Bayesian network probabilistic model(PI Prob). Our model was constructed with clinical-immunology domain expertise, trained on a balanced cohort of 100 cases-controls and validated on an unseen balanced cohort of 150 cases-controls. Performance was measured by area under the receiver operator characteristic curve (AUROC). We also compared our network performance to classic machine learning model performance on the same dataset.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Results</h3><p class="para" id="N65561">PI Prob was accurate in classifying immunodeficiency patients from controls (AUROC = 0.945; p&lt;0.0001) at a risk threshold of ≥6%. Additionally, the model was 89% accurate for categorizing validation cohort members into appropriate International Union of Immunological Societies diagnostic categories. Our network outperformed 3 other machine learning models and provides superior transparency with a prescriptive output element.</p></div><div class="section" id="sec005"><h3 class="BHead" id="nov000-5">Conclusion</h3><p class="para" id="N65567">Artificial intelligence methods can classify risk for primary immunodeficiency and guide management. PI Prob enables accurate, objective decision making about risk and guides the user towards the appropriate diagnostic evaluation for patients with recurrent infections. Probabilistic models can be trained with small datasets underscoring their utility for rare disease detection given appropriate domain expertise for feature selection and network construction.</p></div>]]></description>
            <pubDate><![CDATA[2021-02-16T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[AnnapuRNA: A scoring function for predicting RNA-small molecule binding poses]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765882101158-54ed9d72-3793-485c-863d-1451bb968d64/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pcbi.1008309</link>
            <description><![CDATA[<p class="para" id="N65539">RNA is considered as an attractive target for new small molecule drugs. Designing active compounds can be facilitated by computational modeling. Most of the available tools developed for these prediction purposes, such as molecular docking or scoring functions, are parametrized for protein targets. The performance of these methods, when applied to RNA-ligand systems, is insufficient. To overcome these problems, we developed AnnapuRNA, a new knowledge-based scoring function designed to evaluate RNA-ligand complex structures, generated by any computational docking method. We also evaluated three main factors that may influence the structure prediction, i.e., the starting conformer of a ligand, the docking program, and the scoring function used. We applied the AnnapuRNA method for a <i>post-hoc</i> study of the recently published structures of the FMN riboswitch. Software is available at https://github.com/filipspl/AnnapuRNA.</p><p class="para" id="N65542">Drug development is a lengthy and complicated process, which requires costly experiments on a very large number of chemical compounds. The identification of chemical molecules with desired properties can be facilitated by computational methods. Several methods were developed for computer-aided design of drugs that target protein molecules. However, recently the ribonucleic acid (RNA) emerged as an attractive target for the development of new drugs. Unfortunately, the portfolio of the computer methods that can be applied to study RNA and its interactions with small chemical molecules is very limited. This situation motivated us to develop a new computational method, with which to predict RNA-small molecule interactions. To this end, we collected the information on the statistics of interactions in experimentally determined structures of complexes formed by RNA with small molecules. We then used the statistical data to train machine learning methods aiming to distinguish between RNA-ligand interactions observed experimentally and other interactions that can be observed in theoretical analyses, but are not observed in nature. The resulting method called AnnapuRNA is superior to other similar tools and can be used to predict preferred ligands of RNA molecules and how RNA and small molecules interact with each other.</p>]]></description>
            <pubDate><![CDATA[2021-02-01T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Ten simple rules for engaging with artificial intelligence in biomedicine]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765880515238-843bb75d-768c-4dd7-866e-7b7b8112b91c/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pcbi.1008531</link>
            <description><![CDATA[]]></description>
            <pubDate><![CDATA[2021-02-11T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Multifidelity computing for coupling full and reduced order models]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765878742926-10255191-bbd7-4025-8b78-bd756ee3be8e/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246092</link>
            <description><![CDATA[<p class="para" id="N65539">Hybrid physics-machine learning models are increasingly being used in simulations of transport processes. Many complex multiphysics systems relevant to scientific and engineering applications include multiple spatiotemporal scales and comprise a multifidelity problem sharing an interface between various formulations or heterogeneous computational entities. To this end, we present a robust hybrid analysis and modeling approach combining a physics-based full order model (FOM) and a data-driven reduced order model (ROM) to form the building blocks of an integrated approach among mixed fidelity descriptions toward predictive digital twin technologies. At the interface, we introduce a long short-term memory network to bridge these high and low-fidelity models in various forms of interfacial error correction or prolongation. The proposed interface learning approaches are tested as a new way to address ROM-FOM coupling problems solving nonlinear advection-diffusion flow situations with a bifidelity setup that captures the essence of a broad class of transport processes.</p>]]></description>
            <pubDate><![CDATA[2021-02-11T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Unsupervised learning of Swiss population spatial distribution]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765878381680-ffa0123e-e901-478d-a719-ffeb6440f283/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246529</link>
            <description><![CDATA[<p class="para" id="N65539">The paper deals with the analysis of spatial distribution of Swiss population using fractal concepts and unsupervised learning algorithms. The research methodology is based on the development of a high dimensional feature space by calculating local growth curves, widely used in fractal dimension estimation and on the application of clustering algorithms in order to reveal the patterns of spatial population distribution. The notion “unsupervised” also means, that only some general criteria—density, dimensionality, homogeneity, are used to construct an input feature space, without adding any supervised/expert knowledge. The approach is very powerful and provides a comprehensive local information about density and homogeneity/fractality of spatially distributed point patterns.</p>]]></description>
            <pubDate><![CDATA[2021-02-11T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Construction project risk prediction model based on EW-FAHP and one dimensional convolution neural network]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765877208399-afd8667d-46da-4f15-a1fc-51a9ca2ca408/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246539</link>
            <description><![CDATA[<p class="para" id="N65539">In order to solve the problem of low accuracy of traditional construction project risk prediction, a project risk prediction model based on EW-FAHP and 1D-CNN(One Dimensional Convolution Neural Network) is proposed. Firstly, the risk evaluation index value of construction project is selected by literature analysis method, and the comprehensive weight of risk index is obtained by combining entropy weight method (EW) and fuzzy analytic hierarchy process (FAHP). The risk weight is input into the 1D-CNN model for training and learning, and the prediction values of construction period risk and cost risk are output to realize the risk prediction. The experimental results show that the average absolute error of the construction period risk and cost risk of the risk prediction model proposed in this paper is below 0.1%, which can meet the risk prediction of construction projects with high accuracy.</p>]]></description>
            <pubDate><![CDATA[2021-02-09T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[A pre-training and self-training approach for biomedical named entity recognition]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765873809489-3f6fc036-cade-4c6a-9a44-fd90328b0aab/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246310</link>
            <description><![CDATA[<p class="para" id="N65539">Named entity recognition (NER) is a key component of many scientific literature mining tasks, such as information retrieval, information extraction, and question answering; however, many modern approaches require large amounts of labeled training data in order to be effective. This severely limits the effectiveness of NER models in applications where expert annotations are difficult and expensive to obtain. In this work, we explore the effectiveness of transfer learning and semi-supervised self-training to improve the performance of NER models in biomedical settings with very limited labeled data (250-2000 labeled samples). We first pre-train a BiLSTM-CRF and a BERT model on a very large general biomedical NER corpus such as MedMentions or Semantic Medline, and then we fine-tune the model on a more specific target NER task that has very limited training data; finally, we apply semi-supervised self-training using unlabeled data to further boost model performance. We show that in NER tasks that focus on common biomedical entity types such as those in the Unified Medical Language System (UMLS), combining transfer learning with self-training enables a NER model such as a BiLSTM-CRF or BERT to obtain similar performance with the same model trained on 3x-8x the amount of labeled data. We further show that our approach can also boost performance in a low-resource application where entities types are more rare and not specifically covered in UMLS.</p>]]></description>
            <pubDate><![CDATA[2021-02-09T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[What predicts legislative success of early care and education policies?: Applications of machine learning and Natural Language Processing in a cross-state early childhood policy analysis]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765873733077-cee94f56-6075-414e-8094-a4eb05386959/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246730</link>
            <description><![CDATA[<p class="para" id="N65539">Following the pioneering efforts of a federal Head Start program, U.S. state policymakers have rapidly expanded access to Early Care and Education (ECE) programs with strong bipartisan support. Within the past decade the enrollment of 4 year-olds has roughly doubled in state-funded preschool. Despite these public investments, the content and priorities of early childhood legislation–enacted and failed–have rarely been examined. This study integrates perspectives from public policy, political science, developmental science, and machine learning in examining state ECE bills in identifying key factors associated with legislative success. Drawing from the Early Care and Education Bill Tracking Database, we employed Latent Dirichlet Allocation (LDA), a statistical topic identification model, to examine 2,396 ECE bills across the 50 U.S. states during the 2015-2018. First, a six-topic solution demonstrated the strongest fit theoretically and empirically suggesting two meta policy priorities: ‘ECE finance’ and ‘ECE services’. ‘ECE finance’ comprised three dimensions: (1) Revenues, (2) Expenditures, and (3) Fiscal Governance. ‘ECE services’ also included three dimensions: (1) PreK, (2) Child Care, and (3) Health and Human Services (HHS). Further, we found that bills covering a higher proportion of HHS, Fiscal Governance, or Expenditures were more likely to pass into law relative to bills focusing largely on PreK, Child Care, and Revenues. Additionally, legislative effectiveness of the bill’s primary sponsor was a strong predictor of legislative success, and further moderated the relation between bill content and passage. Highly effective legislators who had previously passed five or more bills had an extremely high probability of introducing a legislation that successfully passed regardless of topic. Legislation with expenditures as policy priorities benefitted the most from having an effective legislator. We conclude with a discussion of the empirical findings within the broader context of early childhood policy literature and suggest implications for future research and policy.</p>]]></description>
            <pubDate><![CDATA[2021-02-11T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Hybrid BW-EDAS MCDM methodology for optimal industrial robot selection]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765873634523-fca0944f-78ae-44fb-b9b5-ba33d7bc9e2a/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246738</link>
            <description><![CDATA[<p class="para" id="N65539">Industrial robots have different capabilities and specifications according to the required applications. It is becoming difficult to select a suitable robot for specific applications and requirements due to the availability of several types with different specifications of robots in the market. Best-worst method is a useful, highly consistent and reliable method to derive weights of criteria and it is worthy to integrate it with the evaluation based on distance from average solution (EDAS) method that is more applicable and needs fewer number of calculations as compared to other methods. An example is presented to show the validity and usability of the proposed methodology. Comparison of ranking results matches with the well-known distance-based approach, technique for order preference by similarity to ideal solution and VIseKriterijumska Optimizacija I Kompromisno Resenje (VIKOR) methods showing the robustness of the best-worst EDAS hybrid method. Sensitivity analysis performed using eighty to one ratio shows that the proposed hybrid MCDM methodology is more stable and reliable.</p>]]></description>
            <pubDate><![CDATA[2021-02-09T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Reference evapotranspiration of Brazil modeled with machine learning techniques and remote sensing]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765873558529-1ad6b14d-1304-4af2-8be5-ab574306f59a/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245834</link>
            <description><![CDATA[<p class="para" id="N65539">Reference evapotranspiration (ETo) is a fundamental parameter for hydrological studies and irrigation management. The Penman-Monteith method is the standard to estimate ETo and requires several meteorological elements. In developing countries, the number of weather stations is insufficient. Thus, free products of remote sensing with evapotranspiration information must be used for this purpose. In this context, the objective of this study was to estimate monthly ETo from potential evapotranspiration (PET) made available by MOD16 product. In this study, the monthly ETo estimated by Penman-Monteith method was considered as the standard. For this, data from 265 weather station of the National Institute of Meteorology (INMET), spread all over the Brazilian territory, were acquired for the period from 2000 to 2014 (15 years). For these months, monthly PET values from MOD16 product for all Brazil were also downloaded. By using machine learning algorithms and information from WorldClim as covariates, ETo was estimated through images from the MOD16 product. To perform the modeling of ETo, eight regression algorithms were tested: multiple linear regression; random forest; cubist; partial least squares; principal components regression; adaptive forward-backward greedy; generalized boosted regression and generalized linear model by likelihood-based boosting. Data from 2000 to 2012 (13 years) were used for training and data of 2013 and 2014 (2 years) were used to test the models. The PET made available by the MOD16 product showed higher values than those of ETo for different periods and climatic regions of Brazil. However, the MOD16 product showed good correlation with ETo, indicating that it can be used in ETo estimation. All models of machine learning were effective in improving the performance of the metrics evaluated. Cubist was the model that presented the best metrics for r<sup>2</sup> (0.91), NSE (0.90) and nRMSE (8.54%) and should be preferred for ETo prediction. MOD16 product is recommended to be used to predict monthly ETo, which opens possibilities for its use in several other studies.</p>]]></description>
            <pubDate><![CDATA[2021-02-09T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Use of artificial intelligence on Electroencephalogram (EEG) waveforms to predict failure in early school grades in children from a rural cohort in Pakistan]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765873453484-f8b1a1f1-367d-4b8d-b991-bad2baf39d8b/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246236</link>
            <description><![CDATA[<p class="para" id="N65539">Universal primary education is critical for individual academic growth and overall adult productivity of nations. Estimates indicate that 25% of 59 million primary age out of school children drop out and early grade failure is one of the factors. An objective and feasible screening measure to identify at-risk children in the early grades can help to design appropriate interventions. The objective of this study was to use a Machine Learning algorithm to evaluate the power of Electroencephalogram (EEG) data collected at age 4 in predicting academic achievement at age 8 among rural children in Pakistan. Demographic and EEG data from 96 children of a cohort along with their academic achievement in grade 1–2 measured using an academic achievement test of Math and language at the age of 7–8 years was used to develop the machine learning algorithm. K- Nearest Neighbor (KNN) classifier was used on different model combinations of EEG, sociodemographic and home environment variables. KNN model was evaluated using 5 Stratified Folds based on the sensitivity and specificity. In the current dataset, 55% and 74% failed in the mathematics and language test respectively. On testing data across each fold, the mean sensitivity and specificity was calculated. Sensitivity was similar when EEG variables were combined with sociodemographic, and home environment (Math = 58.7%, Language = 66.3%) variables but specificity improved (Math = 43.4% to 50.6% and Language = 32% to 60%). The model requires further validation for EEG to be used as a screening measure with adequate sensitivity and specificity to identify children in their preschool age who may be at high risk of failure in early grades.</p>]]></description>
            <pubDate><![CDATA[2021-02-08T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Tracking individual honeybees among wildflower clusters with computer vision-facilitated pollinator monitoring]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765873083608-2bee5012-e281-4e6f-b95e-f21526632088/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0239504</link>
            <description><![CDATA[<p class="para" id="N65539">Monitoring animals in their natural habitat is essential for advancement of animal behavioural studies, especially in pollination studies. Non-invasive techniques are preferred for these purposes as they reduce opportunities for research apparatus to interfere with behaviour. One potentially valuable approach is image-based tracking. However, the complexity of tracking unmarked wild animals using video is challenging in uncontrolled outdoor environments. Out-of-the-box algorithms currently present several problems in this context that can compromise accuracy, especially in cases of occlusion in a 3D environment. To address the issue, we present a novel hybrid detection and tracking algorithm to monitor unmarked insects outdoors. Our software can detect an insect, identify when a tracked insect becomes occluded from view and when it re-emerges, determine when an insect exits the camera field of view, and our software assembles a series of insect locations into a coherent trajectory. The insect detecting component of the software uses background subtraction and deep learning-based detection together to accurately and efficiently locate the insect among a cluster of wildflowers. We applied our method to track honeybees foraging outdoors using a new dataset that includes complex background detail, wind-blown foliage, and insects moving into and out of occlusion beneath leaves and among three-dimensional plant structures. We evaluated our software against human observations and previous techniques. It tracked honeybees at a rate of 86.6% on our dataset, 43% higher than the computationally more expensive, standalone deep learning model YOLOv2. We illustrate the value of our approach to quantify fine-scale foraging of honeybees. The ability to track unmarked insect pollinators in this way will help researchers better understand pollination ecology. The increased efficiency of our hybrid approach paves the way for the application of deep learning-based techniques to animal tracking in real-time using low-powered devices suitable for continuous monitoring.</p>]]></description>
            <pubDate><![CDATA[2021-02-11T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Predicting the rental value of houses in household surveys in Tanzania, Uganda and Malawi: Evaluations of hedonic pricing and machine learning approaches]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765873061330-2018f824-6658-4de5-a2e0-2f8d0d021d10/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244953</link>
            <description><![CDATA[<p class="para" id="N65539">Housing value is a major component of the aggregate expenditure used in the analyses of welfare status of households in the development economics literature. Therefore, an accurate estimation of housing services is important to obtain the value of housing in household surveys. Data show that a significant proportion of households in a typical Living Standard Measurement Survey (LSMS), adopted by the Word Bank and others, are self-owned. The standard approach to predict the housing value for such surveys is based on the rental cost of the house. A hedonic pricing applying an Ordinary Least Squares (OLS) method is normally used to predict rental values. The literature shows that Machine Learning (ML) methods, shown to uncover generalizable patterns based on a given data, have better predictive power over OLS applied in other valuation exercises. We examined whether or not a class of ML methods (e.g. Ridge, LASSO, Tree, Bagging, Random Forest, and Boosting) provided superior prediction of rental value of housing over OLS methods accounting for spatial autocorrelations using household level survey data from Uganda, Tanzania, and Malawi, across multiple years. Our results showed that the Machine Learning methods (Boosting, Bagging, Forest, Ridge and LASSO) are the best models in predicting house values using out-of-sample data set for all the countries and all the years. On the other hand, Tree regression underperformed relative to the various OLS models, over the same data sets. With the availability of abundant data and better computing power, ML methods provide viable alternative to predicting housing values in household surveys.</p>]]></description>
            <pubDate><![CDATA[2021-02-11T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Data structure <i>set-trie</i> for storing and querying sets: Theoretical and empirical analysis]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765872704212-c00c3490-8d09-4c9d-b94c-cce76532a5e3/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245122</link>
            <description><![CDATA[<p class="para" id="N65539">Set containment operations form an important tool in various fields such as information retrieval, AI systems, object-relational databases, and Internet applications. In the paper, a <i>set-trie</i> data structure for storing sets is considered, along with the efficient algorithms for the corresponding set containment operations. We present the mathematical and empirical study of the set-trie. In the mathematical study, the relevant upper-bounds on the efficiency of its expected performance are established by utilizing a natural probabilistic model. In the empirical study, we give insight into how different distributions of input data impact the efficiency of set-trie. Using the correct parameters for those randomly generated datasets, we expose the key sources of the input sensitivity of set-trie. Finally, the empirical comparison of set-trie with the inverted index is based on the real-world datasets containing sets of low cardinality. The comparison shows that the running time of set-trie consistently outperforms the inverted index by orders of magnitude.</p>]]></description>
            <pubDate><![CDATA[2021-02-10T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Using machine learning to investigate the relationship between domains of functioning and functional mobility in older adults]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765872474016-bb256252-ec0e-4fa6-8850-97a6636ff1d8/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246397</link>
            <description><![CDATA[<p class="para" id="N65539">Previous studies have shown that functional mobility, along with other physical functions, decreases with advanced age. However, it is still unclear which domains of functioning (body structures, body functions, and activities) are most closely related to functional mobility. This study used machine learning classification to predict the rankings of Timed Up and Go tests based on the results of four assessments (soft lean mass, FEV<sub>1</sub>/FVC, knee extension torque, and one-leg standing time). We tested whether assessment results for each level could predict functional mobility assessments in older adults. Using support vector machines for machine learning classification, we verified that the four assessments of each level could classify functional mobility. Knee extension torque (from the body function domain) was the most closely related assessment. Naturally, the classification accuracy rate increased with a larger number of assessments as explanatory variables. However, knee extension torque remained the highest of all assessments. This extended to all combinations (of 2–3 assessments) that included knee extension torque. This suggests that resistance training may help protect individuals suffering from age-related declines in functional mobility.</p>]]></description>
            <pubDate><![CDATA[2021-02-11T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[A deep learning-based method for grip strength prediction: Comparison of multilayer perceptron and polynomial regression approaches]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765872385503-8941cfef-76bd-47c9-b73b-2b678ee87bba/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246870</link>
            <description><![CDATA[<p class="para" id="N65539">The objective of this study was to accurately predict the grip strength using a deep learning-based method (e.g., multi-layer perceptron [MLP] regression). The maximal grip strength with varying postures (upper arm, forearm, and lower body) of 164 young adults (100 males and 64 females) were collected. The data set was divided into a training set (90% of data) and a test set (10% of data). Different combinations of variables including demographic and anthropometric information of individual participants and postures was tested and compared to find the most predictive model. The MLP regression and 3 different polynomial regressions (linear, quadratic, and cubic) were conducted and the performance of regression was compared. The results showed that including all variables showed better performance than other combinations of variables. In general, MLP regression showed higher performance than polynomial regressions. Especially, MLP regression considering all variables achieved the highest performance of grip strength prediction (RMSE = 69.01N, R = 0.88, ICC = 0.92). This deep learning-based regression (MLP) would be useful to predict on-site- and individual-specific grip strength in the workspace to reduce the risk of musculoskeletal disorders in the upper extremity.</p>]]></description>
            <pubDate><![CDATA[2021-02-11T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[A novel cross-validation strategy for artificial neural networks using distributed-lag environmental factors]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765854758897-c31ea49c-477c-476f-84f9-5a9a34f2afcb/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244094</link>
            <description><![CDATA[<p class="para" id="N65539">In recent years, machine learning methods have been applied to various prediction scenarios in time-series data. However, some processing procedures such as cross-validation (CV) that rearrange the order of the longitudinal data might ruin the seriality and lead to a potentially biased outcome. Regarding this issue, a recent study investigated how different types of CV methods influence the predictive errors in conventional time-series data. Here, we examine a more complex distributed lag nonlinear model (DLNM), which has been widely used to assess the cumulative impacts of past exposures on the current health outcome. This research extends the DLNM into an artificial neural network (ANN) and investigates how the ANN model reacts to various CV schemes that result in different predictive biases. We also propose a newly designed permutation ratio to evaluate the performance of the CV in the ANN. This ratio mimics the concept of the R-square in conventional statistical regression models. The results show that as the complexity of the ANN increases, the predicted outcome becomes more stable, and the bias shows a decreasing trend. Among the different settings of hyperparameters, the novel strategy, Leave One Block Out Cross-Validation (LOBO-CV), demonstrated much better results, and the lowest mean square error was observed. The hyperparameters of the ANN trained by the LOBO-CV yielded the minimum number of prediction errors. The newly proposed permutation ratio indicates that LOBO-CV can contribute up to 34% of the prediction accuracy.</p>]]></description>
            <pubDate><![CDATA[2021-01-07T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[A framework for the risk prediction of avian influenza occurrence: An Indonesian case study]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765854025713-ab498abe-490a-4189-a220-6c949af999c2/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245116</link>
            <description><![CDATA[<p class="para" id="N65539">Avian influenza viruses can cause economically devastating diseases in poultry and have the potential for zoonotic transmission. To mitigate the consequences of avian influenza, disease prediction systems have become increasingly important. In this study, we have proposed a framework for the prediction of the occurrence and spread of avian influenza events in a geographical area. The application of the proposed framework was examined in an Indonesian case study. An extensive list of historical data sources containing disease predictors and target variables was used to build spatiotemporal and transactional datasets. To combine disparate sources, data rows were scaled to a temporal scale of 1-week and a spatial scale of 1-degree × 1-degree cells. Given the constructed datasets, underlying patterns in the form of rules explaining the risk of occurrence and spread of avian influenza were discovered. The created rules were combined and ordered based on their importance and then stored in a knowledge base. The results suggested that the proposed framework could act as a tool to gain a broad understanding of the drivers of avian influenza epidemics and may facilitate the prediction of future disease events.</p>]]></description>
            <pubDate><![CDATA[2021-01-15T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Automatic classification of mice vocalizations using Machine Learning techniques and Convolutional Neural Networks]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765853504160-f85ec4d7-73d9-4cc8-83a9-78e323467636/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244636</link>
            <description><![CDATA[<p class="para" id="N65539">Ultrasonic vocalizations (USVs) analysis is a well-recognized tool to investigate animal communication. It can be used for behavioral phenotyping of murine models of different disorders. The USVs are usually recorded with a microphone sensitive to ultrasound frequencies and they are analyzed by specific software. Different calls typologies exist, and each ultrasonic call can be manually classified, but the qualitative analysis is highly time-consuming. Considering this framework, in this work we proposed and evaluated a set of supervised learning methods for automatic USVs classification. This could represent a sustainable procedure to deeply analyze the ultrasonic communication, other than a standardized analysis. We used manually built datasets obtained by segmenting the USVs audio tracks analyzed with the Avisoft software, and then by labelling each of them into 10 representative classes. For the automatic classification task, we designed a Convolutional Neural Network that was trained receiving as input the spectrogram images associated to the segmented audio files. In addition, we also tested some other supervised learning algorithms, such as Support Vector Machine, Random Forest and Multilayer Perceptrons, exploiting informative numerical features extracted from the spectrograms. The performance showed how considering the whole time/frequency information of the spectrogram leads to significantly higher performance than considering a subset of numerical features. In the authors’ opinion, the experimental results may represent a valuable benchmark for future work in this research field.</p>]]></description>
            <pubDate><![CDATA[2021-01-19T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[TEM image restoration from fast image streams]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765852934781-c3bf7141-f66a-4ebe-a3a6-61314d41c8a9/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246336</link>
            <description><![CDATA[<p class="para" id="N65539">Microscopy imaging experiments generate vast amounts of data, and there is a high demand for smart acquisition and analysis methods. This is especially true for transmission electron microscopy (TEM) where terabytes of data are produced if imaging a full sample at high resolution, and analysis can take several hours. One way to tackle this issue is to collect a continuous stream of low resolution images whilst moving the sample under the microscope, and thereafter use this data to find the parts of the sample deemed most valuable for high-resolution imaging. However, such image streams are degraded by both motion blur and noise. Building on deep learning based approaches developed for deblurring videos of natural scenes we explore the opportunities and limitations of deblurring and denoising images captured from a fast image stream collected by a TEM microscope. We start from existing neural network architectures and make adjustments of convolution blocks and loss functions to better fit TEM data. We present deblurring results on two real datasets of images of kidney tissue and a calibration grid. Both datasets consist of low quality images from a fast image stream captured by moving the sample under the microscope, and the corresponding high quality images of the same region, captured after stopping the movement at each position to let all motion settle. We also explore the generalizability and overfitting on real and synthetically generated data. The quality of the restored images, evaluated both quantitatively and visually, show that using deep learning for image restoration of TEM live image streams has great potential but also comes with some limitations.</p>]]></description>
            <pubDate><![CDATA[2021-02-01T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Deep learning framework for subject-independent emotion detection using wireless signals]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765852053621-5e5f58b3-3849-487a-b034-e0b01f137b4d/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0242946</link>
            <description><![CDATA[<p class="para" id="N65539">Emotion states recognition using wireless signals is an emerging area of research that has an impact on neuroscientific studies of human behaviour and well-being monitoring. Currently, standoff emotion detection is mostly reliant on the analysis of facial expressions and/or eye movements acquired from optical or video cameras. Meanwhile, although they have been widely accepted for recognizing human emotions from the multimodal data, machine learning approaches have been mostly restricted to subject dependent analyses which lack of generality. In this paper, we report an experimental study which collects heartbeat and breathing signals of 15 participants from radio frequency (RF) reflections off the body followed by novel noise filtering techniques. We propose a novel deep neural network (DNN) architecture based on the fusion of raw RF data and the processed RF signal for classifying and visualising various emotion states. The proposed model achieves high classification accuracy of 71.67% for independent subjects with 0.71, 0.72 and 0.71 precision, recall and F1-score values respectively. We have compared our results with those obtained from five different classical ML algorithms and it is established that deep learning offers a superior performance even with limited amount of raw RF and post processed time-sequence data. The deep learning model has also been validated by comparing our results with those from ECG signals. Our results indicate that using wireless signals for stand-by emotion state detection is a better alternative to other technologies with high accuracy and have much wider applications in future studies of behavioural sciences.</p>]]></description>
            <pubDate><![CDATA[2021-02-03T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[3D multi-scale deep convolutional neural networks for pulmonary nodule detection]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765851921201-1718d049-57a1-458d-8451-7dcb2ea501b1/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244406</link>
            <description><![CDATA[<p class="para" id="N65539">With the rapid development of big data and artificial intelligence technology, computer-aided pulmonary nodule detection based on deep learning has achieved some successes. However, the sizes of pulmonary nodules vary greatly, and the pulmonary nodules have visual similarity with structures such as blood vessels and shadows around pulmonary nodules, which make the quick and accurate detection of pulmonary nodules in CT image still a challenging task. In this paper, we propose two kinds of 3D multi-scale deep convolution neural networks for nodule candidate detection and false positive reduction respectively. Among them, the nodule candidate detection network consists of two parts: 1) the backbone network part Res2SENet, which is used to extract multi-scale feature information of pulmonary nodules, it is composed of the multi-scale Res2Net modules of multiple available receptive fields at a granular level and the squeeze-and-excitation units; 2) the detection part, which uses a region proposal network structure to determine region candidates, and introduces context enhancement module and spatial attention module to improve detection performance. The false positive reduction network, also composed of the multi-scale Res2Net modules and the squeeze-and-excitation units, can further classify the nodule candidates generated by the nodule candidate detection network and screen out the ground truth positive nodules. Finally, the prediction probability generated by the nodule candidate detection network is weighted average with the prediction probability generated by the false positive reduction network to obtain the final results. The experimental results on the publicly available LUNA16 dataset showed that the proposed method has a superior ability to detect pulmonary nodules in CT images.</p>]]></description>
            <pubDate><![CDATA[2021-01-07T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Validity and precision of the International Physical Activity Questionnaire for climacteric women using computational intelligence techniques]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765851583563-35ad0564-8ce1-440d-9f95-3ea485eb8abf/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245240</link>
            <description><![CDATA[<p class="para" id="N65539">This study aimed to evaluate the validity and precision of the International Physical Activity Questionnaire (IPAQ) for climacteric women using computational intelligence techniques. The instrument was applied to 873 women aged between 40 and 65 years. Considering the proposal to regroup the set of data related to the level of physical activity of climacteric women using the IPAQ, we used 2 algorithms: Kohonen and k-means, and, to evaluate the validity of these clusters, 3 indexes were used: Silhouette, PBM and Dunn. The questionnaire was tested for validity (factor analysis) and precision (Cronbach's alpha). The Random Forests technique was used to assess the importance of the variables that make up the IPAQ. To classify these variables, we used 3 algorithms: Suport Vector Machine, Artificial Neural Network and Decision Tree. The results of the tests to evaluate the clusters suggested that what is recommended for IPAQ, when applied to climacteric women, is to categorize the results into two groups. The factor analysis resulted in three factors, with factor 1 being composed of variables 3 to 6; factor 2 for variables 7 and 8; and factor 3 for variables 1 and 2. Regarding the reliability estimate, the results of the standardized Cronbach's alpha test showed values between 0.63 to 0.85, being considered acceptable for the construction of the construct. In the test of importance of the variables that make up the instrument, the results showed that variables 1 and 8 presented a lesser degree of importance and by the analysis of Accuracy, Recall, Precision and area under the ROC curve, there was no variation when the results were analyzed with all IPAQ variables but variables 1 and 8. Through this analysis, we concluded that the IPAQ, short version, has adequate measurement properties for the investigated population.</p>]]></description>
            <pubDate><![CDATA[2021-01-14T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Novel online Recommendation algorithm for Massive Open Online Courses (NoR-MOOCs)]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765850537949-bb54554d-4d8c-453e-9796-657d4c7d7104/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245485</link>
            <description><![CDATA[<p class="para" id="N65539">Massive Open Online Courses (MOOCs) have gained in popularity over the last few years. The space of online learning resources has been increasing exponentially and has created a problem of information overload. To overcome this problem, recommender systems that can recommend learning resources to users according to their interests have been proposed. MOOCs contain a huge amount of data with the quantity of data increasing as new learners register. Traditional recommendation techniques suffer from scalability, sparsity and cold start problems resulting in poor quality recommendations. Furthermore, they cannot accommodate the incremental update of the model with the arrival of new data making them unsuitable for MOOCs dynamic environment. From this line of research, we propose a novel online recommender system, namely NoR-MOOCs, that is accurate, scales well with the data and moreover overcomes previously recorded problems with recommender systems. Through extensive experiments conducted over the COCO data-set, we have shown empirically that NoR-MOOCs significantly outperforms traditional KMeans and Collaborative Filtering algorithms in terms of predictive and classification accuracy metrics.</p>]]></description>
            <pubDate><![CDATA[2021-01-22T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Automatic clustering method to segment COVID-19 CT images]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765850404279-93101a20-e8ad-4e2c-981d-a418d1171aa4/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244416</link>
            <description><![CDATA[<p class="para" id="N65539">Coronavirus pandemic (COVID-19) has infected more than ten million persons worldwide. Therefore, researchers are trying to address various aspects that may help in diagnosis this pneumonia. Image segmentation is a necessary pr-processing step that implemented in image analysis and classification applications. Therefore, in this study, our goal is to present an efficient image segmentation method for COVID-19 Computed Tomography (CT) images. The proposed image segmentation method depends on improving the density peaks clustering (DPC) using generalized extreme value (GEV) distribution. The DPC is faster than other clustering methods, and it provides more stable results. However, it is difficult to determine the optimal number of clustering centers automatically without visualization. So, GEV is used to determine the suitable threshold value to find the optimal number of clustering centers that lead to improving the segmentation process. The proposed model is applied for a set of twelve COVID-19 CT images. Also, it was compared with traditional k-means and DPC algorithms, and it has better performance using several measures, such as PSNR, SSIM, and Entropy.</p>]]></description>
            <pubDate><![CDATA[2021-01-08T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Machine learning model for predicting severity prognosis in patients infected with COVID-19: Study protocol from COVID-AI Brasil]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765850237260-57e50356-4de9-419c-8685-e3dddb5f3dcf/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245384</link>
            <description><![CDATA[<p class="para" id="N65539">The new coronavirus, which began to be called SARS-CoV-2, is a single-stranded RNA beta coronavirus, initially identified in Wuhan (Hubei province, China) and currently spreading across six continents causing a considerable harm to patients, with no specific tools until now to provide prognostic outcomes. Thus, the aim of this study is to evaluate possible findings on chest CT of patients with signs and symptoms of respiratory syndromes and positive epidemiological factors for COVID-19 infection and to correlate them with the course of the disease. In this sense, it is also expected to develop specific machine learning algorithm for this purpose, through pulmonary segmentation, which can predict possible prognostic factors, through more accurate results. Our alternative hypothesis is that the machine learning model based on clinical, radiological and epidemiological data will be able to predict the severity prognosis of patients infected with COVID-19. We will perform a multicenter retrospective longitudinal study to obtain a large number of cases in a short period of time, for better study validation. Our convenience sample (at least 20 cases for each outcome) will be collected in each center considering the inclusion and exclusion criteria. We will evaluate patients who enter the hospital with clinical signs and symptoms of acute respiratory syndrome, from March to May 2020. We will include individuals with signs and symptoms of acute respiratory syndrome, with positive epidemiological history for COVID-19, who have performed a chest computed tomography. We will assess chest CT of these patients and to correlate them with the course of the disease. Primary outcomes:1) Time to hospital discharge; 2) Length of stay in the ICU; 3) orotracheal intubation;4) Development of Acute Respiratory Discomfort Syndrome. Secondary outcomes:1) Sepsis; 2) Hypotension or cardiocirculatory dysfunction requiring the prescription of vasopressors or inotropes; 3) Coagulopathy; 4) Acute Myocardial Infarction; 5) Acute Renal Insufficiency; 6) Death. We will use the AUC and F1-score of these algorithms as the main metrics, and we hope to identify algorithms capable of generalizing their results for each specified primary and secondary outcome.</p>]]></description>
            <pubDate><![CDATA[2021-02-01T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Improving precise point positioning performance based on Prophet model]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765848376165-12e1167d-62e5-4991-a4d6-aef2d77167e0/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245561</link>
            <description><![CDATA[<p class="para" id="N65539">Precision point positioning (PPP) is widely used in maritime navigation and other scenarios because it does not require a reference station. In PPP, the satellite clock bias (SCB) cannot be eliminated by differential, thus leading to an increase in positioning error. The prediction accuracy of SCB has become one of the key factors restricting positioning accuracy. Although International GNSS Service (IGS) provides the ultra-rapid ephemeris prediction part (IGU-P), its quality and real-time performance can not meet the practical application. In order to improve the accuracy of PPP, this paper proposes to use the Prophet model to predict SCB. Specifically, SCB sequence is read from the observation part in the ultra-rapid ephemeris (IGU-O) released by IGS. Next, the SCB sequence between adjacent epochs are subtracted to obtain the corresponding SCB single difference sequence. Then using the Prophet model to predict SCB single difference sequence. Finally, the prediction result is substituted into the PPP positioning observation equation to obtain the positioning result. This paper uses the final ephemeris (IGF) published by IGS as a benchmark and compares the experimental results with IGU-P. For the selected four satellites, compared with the results of the IGU-P, the accuracy of SCB prediction of the model in this paper is improved by about 50.3%, 61.7%, 60.4%, and 48.8%. In terms of PPP positioning results, we use Real-time kinematic (RTK) measurements as a benchmark in this paper. Positioning accuracy has increased by 26%, 35%, and 19% in the N, E, and U directions, respectively. The results show that the Prophet model can improve the performance of PPP.</p>]]></description>
            <pubDate><![CDATA[2021-01-19T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Sustainability planning in the US response to the opioid crisis: An examination using expert and text mining approaches]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765848243455-87b59cbb-4515-4d70-86b9-478aaaf9573a/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245920</link>
            <description><![CDATA[<p class="para" id="N65539">Between January 2016 and June 2020, the Substance Abuse and Mental Health Services Administration rapidly distributed $7.5 billion in response to the U.S. opioid crisis. These funds are designed to increase access to medications for addiction treatment, reduce unmet treatment need, reduce overdose death rates, and provide and sustain effective prevention, treatment and recovery activities. It is unclear whether or not the services developed using these funds will be sustained beyond the start-up period. Based on 34 (64%) State Opioid Response (SOR) applications, we assessed the states’ sustainability plans focusing on potential funding sources, policies, and quality monitoring. We found variable commitment to sustainability across response plans with less than half the states adequately describing sustainability plans. States with higher proportions of opioid prescribing, opioid misuse, and poverty had somewhat higher scores on sustainment. A text mining/machine learning approach automatically rated sustainability in SOR applications with an 82% accuracy compared to human ratings. Because life saving evidence-based programs and services may be lost, intentional commitment to sustainment beyond the bolus of start-up funding is essential.</p>]]></description>
            <pubDate><![CDATA[2021-01-28T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Identification of drug combinations on the basis of machine learning to maximize anti-aging effects]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765848034422-9d071a2c-e107-4c3c-ad10-0a9db5b891aa/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246106</link>
            <description><![CDATA[<p class="para" id="N65539">Aging is a multifactorial process that involves numerous genetic changes, so identifying anti-aging agents is quite challenging. Age-associated genetic factors must be better understood to search appropriately for anti-aging agents. We utilized an aging-related gene expression pattern-trained machine learning system that can implement reversible changes in aging by linking combinatory drugs. <i>In silico</i> gene expression pattern-based drug repositioning strategies, such as connectivity map, have been developed as a method for unique drug discovery. However, these strategies have limitations such as lists that differ for input and drug-inducing genes or constraints to compare experimental cell lines to target diseases. To address this issue and improve the prediction success rate, we modified the original version of expression profiles with a stepwise-filtered method. We utilized a machine learning system called deep-neural network (DNN). Here we report that combinational drug pairs using differential expressed genes (DEG) had a more enhanced anti-aging effect compared with single independent treatments on leukemia cells. This study shows potential drug combinations to retard the effects of aging with higher efficacy using innovative machine learning techniques.</p>]]></description>
            <pubDate><![CDATA[2021-01-28T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Kinematics and workspace analysis of 4SPRR-SPR parallel robots]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765847540212-b052a0fa-6552-4489-9969-41b6e7cbb52e/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0239150</link>
            <description><![CDATA[<p class="para" id="N65539">The 4SPRR-SPR parallel robot, which has considerable potential for application in the field of machining, is a novel closed-loop mechanism with a high rigid-weight ratio. Kinematics and workspace analyses of the 4SPRR-SPR parallel robot are key requirements for its application in machining. In this study, the inverse kinematics of the 4SPRR-SPR parallel robot is analyzed using a geometric method based on the mechanism arrangement of the robot. The forward kinematics model is derived by training the vector-quantified temporal associative memory (VQTAM) network, which originates from a self-organizing map (SOM). Furthermore, an improved algorithm is obtained by combining the locally linear embedding (LLE) and VQTAM methods. A boundary extraction algorithm for the workspace analysis of the parallel robot is proposed. The performance of the boundary extraction algorithm is analyzed and compared with that of a global search algorithm; the result indicates that the novel algorithm has the same computational accuracy in addition to higher efficiency. The workspace of the 4SPRR-SPR parallel robot is analyzed using the boundary extraction algorithm. Finally, the 3D model of the 4SPRR-SPR parallel robot is simulated using the ADAMS software to verify the reliability of the proposed algorithms. The simulation results demonstrate the effectiveness of the methods proposed in this study. In addition, the robot kinematics and workspace analysis methods described herein can be extended to other serial and parallel robots. This research provides a theoretical framework for trajectory planning of mechanisms, workspace optimization of robots, and robotic control.</p>]]></description>
            <pubDate><![CDATA[2021-01-20T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Machine learning-based prediction of in-hospital mortality using admission laboratory data: A retrospective, single-site study using electronic health record data]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765847486228-b9b18dd7-025e-4316-9f2e-b891565315cf/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246640</link>
            <description><![CDATA[<p class="para" id="N65539">Risk assessment of in-hospital mortality of patients at the time of hospitalization is necessary for determining the scale of required medical resources for the patient depending on the patient’s severity. Because recent machine learning application in the clinical area has been shown to enhance prediction ability, applying this technique to this issue can lead to an accurate prediction model for in-hospital mortality prediction. In this study, we aimed to generate an accurate prediction model of in-hospital mortality using machine learning techniques. Patients 18 years of age or older admitted to the University of Tokyo Hospital between January 1, 2009 and December 26, 2017 were used in this study. The data were divided into a training/validation data set (n = 119,160) and a test data set (n = 33,970) according to the time of admission. The prediction target of the model was the in-hospital mortality within 14 days. To generate the prediction model, 25 variables (age, sex, 21 laboratory test items, length of stay, and mortality) were used to predict in-hospital mortality. Logistic regression, random forests, multilayer perceptron, and gradient boost decision trees were performed to generate the prediction models. To evaluate the prediction capability of the model, the model was tested using a test data set. Mean probabilities obtained from trained models with five-fold cross-validation were used to calculate the area under the receiver operating characteristic (AUROC) curve. In a test stage using the test data set, prediction models of in-hospital mortality within 14 days showed AUROC values of 0.936, 0.942, 0.942, and 0.938 for logistic regression, random forests, multilayer perceptron, and gradient boosting decision trees, respectively. Machine learning-based prediction of short-term in-hospital mortality using admission laboratory data showed outstanding prediction capability and, therefore, has the potential to be useful for the risk assessment of patients at the time of hospitalization.</p>]]></description>
            <pubDate><![CDATA[2021-02-05T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Application of deep learning algorithm to detect and visualize vertebral fractures on plain frontal radiographs]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765847454789-8f179b1b-c666-4c4e-9447-e76e0272a3c4/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245992</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Background</h3><p class="para" id="N65543">Identification of vertebral fractures (VFs) is critical for effective secondary fracture prevention owing to their association with the increasing risks of future fractures. Plain abdominal frontal radiographs (PARs) are a common investigation method performed for a variety of clinical indications and provide an ideal platform for the opportunistic identification of VF. This study uses a deep convolutional neural network (DCNN) to identify the feasibility for the screening, detection, and localization of VFs using PARs.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Methods</h3><p class="para" id="N65549">A DCNN was pretrained using ImageNet and retrained with 1306 images from the PARs database obtained between August 2015 and December 2018. The accuracy, sensitivity, specificity, and area under the receiver operating characteristic curve (AUC) were evaluated. The visualization algorithm gradient-weighted class activation mapping (Grad-CAM) was used for model interpretation.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Results</h3><p class="para" id="N65555">Only 46.6% (204/438) of the VFs were diagnosed in the original PARs reports. The algorithm achieved 73.59% accuracy, 73.81% sensitivity, 73.02% specificity, and an AUC of 0.72 in the VF identification.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Conclusion</h3><p class="para" id="N65561">Computer driven solutions integrated with the DCNN have the potential to identify VFs with good accuracy when used opportunistically on PARs taken for a variety of clinical purposes. The proposed model can help clinicians become more efficient and economical in the current clinical pathway of fragile fracture treatment.</p></div>]]></description>
            <pubDate><![CDATA[2021-01-28T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Enhancing web search result clustering model based on multiview multirepresentation consensus cluster ensemble (mmcc) approach]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765846762827-0a9e7fb0-1fcc-4374-ba9c-365446247f1b/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245264</link>
            <description><![CDATA[<p class="para" id="N65539">Existing text clustering methods utilize only one representation at a time (single view), whereas multiple views can represent documents. The multiview multirepresentation method enhances clustering quality. Moreover, existing clustering methods that utilize more than one representation at a time (multiview) use representation with the same nature. Hence, using multiple views that represent data in a different representation with clustering methods is reasonable to create a diverse set of candidate clustering solutions. On this basis, an effective dynamic clustering method must consider combining multiple views of data including semantic view, lexical view (word weighting), and topic view as well as the number of clusters. The main goal of this study is to develop a new method that can improve the performance of web search result clustering (WSRC). An enhanced multiview multirepresentation consensus clustering ensemble (MMCC) method is proposed to create a set of diverse candidate solutions and select a high-quality overlapping cluster. The overlapping clusters are obtained from the candidate solutions created by different clustering methods. The framework to develop the proposed MMCC includes numerous stages: (1) acquiring the standard datasets (MORESQUE and Open Directory Project-239), which are used to validate search result clustering algorithms, (2) preprocessing the dataset, (3) applying multiview multirepresentation clustering models, (4) using the radius-based cluster number estimation algorithm, and (5) employing the consensus clustering ensemble method. Results show an improvement in clustering methods when multiview multirepresentation is used. More importantly, the proposed MMCC model improves the overall performance of WSRC compared with all single-view clustering models.</p>]]></description>
            <pubDate><![CDATA[2021-01-15T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Optimizing multi-supplier multi-item joint replenishment problem for non-instantaneous deteriorating items with quantity discounts]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765841675888-5641545c-445e-4f9c-aa46-c321732d218e/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246035</link>
            <description><![CDATA[<p class="para" id="N65539">This paper deals with a new joint replenishment problem, in which a number of non-instantaneous deteriorating items are replenished from several suppliers under different quantity discounts schemes. Involving both joint replenishment decisions and supplier selection decisions makes the problem to be NP-hard. In particular, the consideration of non-instantaneous deterioration makes it more challenging to handle. We first construct a mathematical model integrated with a supplier selection system and a joint replenishment program for non-instantaneous deteriorating items to formulate the problem. Then we develop a novel swarm intelligence optimization algorithm, the Improved Moth-flame Optimization (IMFO) algorithm, to solve the proposed model. The results of several numerical experiments analyses reveal that the IMFO algorithm is an effective algorithm for solving the proposed model in terms of solution quality and searching stableness. Finally, we conduct extensive experiments to further investigate the performance of the proposed model.</p>]]></description>
            <pubDate><![CDATA[2021-02-08T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Addressless: A new internet server model to prevent network scanning]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765841432270-001873a2-7fd6-4426-a0f8-d527e0dea846/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246293</link>
            <description><![CDATA[<p class="para" id="N65539">Eliminating unnecessary exposure is a principle of server security. The huge IPv6 address space enhances security by making scanning infeasible, however, with recent advances of IPv6 scanning technologies, network scanning is again threatening server security. In this paper, we propose a new model named addressless server, which separates the server into an entrance module and a main service module, and assigns an IPv6 prefix instead of an IPv6 address to the main service module. The entrance module generates a legitimate IPv6 address under this prefix by encrypting the client address, so that the client can access the main server on a destination address that is different in each connection. In this way, the model provides isolation to the main server, prevents network scanning, and minimizes exposure. Moreover it provides a novel framework that supports flexible load balancing, high-availability, and other desirable features. The model is simple and does not require any modification to the client or the network. We implement a prototype and experiments show that our model can prevent the main server from being scanned at a slight performance cost.</p>]]></description>
            <pubDate><![CDATA[2021-02-02T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Mediating artificial intelligence developments through negative and positive incentives]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765841154195-b24dbef8-c78c-4cb7-8189-2c752614f927/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244592</link>
            <description><![CDATA[<p class="para" id="N65539">The field of Artificial Intelligence (AI) is going through a period of great expectations, introducing a certain level of anxiety in research, business and also policy. This anxiety is further energised by an AI race narrative that makes people believe they might be missing out. Whether real or not, a belief in this narrative may be detrimental as some stake-holders will feel obliged to cut corners on safety precautions, or ignore societal consequences just to “win”. Starting from a baseline model that describes a broad class of technology races where winners draw a significant benefit compared to others (such as AI advances, patent race, pharmaceutical technologies), we investigate here how positive (rewards) and negative (punishments) incentives may beneficially influence the outcomes. We uncover conditions in which punishment is either capable of reducing the development speed of unsafe participants or has the capacity to reduce innovation through over-regulation. Alternatively, we show that, in several scenarios, rewarding those that follow safety measures may increase the development speed while ensuring safe choices. Moreover, in the latter regimes, rewards do not suffer from the issue of over-regulation as is the case for punishment. Overall, our findings provide valuable insights into the nature and kinds of regulatory actions most suitable to improve safety compliance in the contexts of both smooth and sudden technological shifts.</p>]]></description>
            <pubDate><![CDATA[2021-01-26T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[HDG-select: A novel GUI based application for gene selection and classification in high dimensional datasets]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765841055869-7021f49c-900c-424e-b362-85e9c6358330/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246039</link>
            <description><![CDATA[<p class="para" id="N65539">The selection and classification of genes is essential for the identification of related genes to a specific disease. Developing a user-friendly application with combined statistical rigor and machine learning functionality to help the biomedical researchers and end users is of great importance. In this work, a novel stand-alone application, which is based on graphical user interface (GUI), is developed to perform the full functionality of gene selection and classification in high dimensional datasets. The so-called HDG-select application is validated on eleven high dimensional datasets of the format CSV and GEO soft. The proposed tool uses the efficient algorithm of combined filter-GBPSO-SVM and it was made freely available to users. It was found that the proposed HDG-select outperformed other tools reported in literature and presented a competitive performance, accessibility, and functionality.</p>]]></description>
            <pubDate><![CDATA[2021-01-28T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Variational-LSTM autoencoder to forecast the spread of coronavirus across the globe]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765840807809-cd1c259d-938a-48d9-9d6f-fe36da871f2e/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246120</link>
            <description><![CDATA[<p class="para" id="N65539">Modelling the spread of coronavirus globally while learning trends at global and country levels remains crucial for tackling the pandemic. We introduce a novel variational-LSTM Autoencoder model to predict the spread of coronavirus for each country across the globe. This deep Spatio-temporal model does not only rely on historical data of the virus spread but also includes factors related to urban characteristics represented in locational and demographic data (such as population density, urban population, and fertility rate), an index that represents the governmental measures and response amid toward mitigating the outbreak (includes 13 measures such as: 1) school closing, 2) workplace closing, 3) cancelling public events, 4) close public transport, 5) public information campaigns, 6) restrictions on internal movements, 7) international travel controls, 8) fiscal measures, 9) monetary measures, 10) emergency investment in health care, 11) investment in vaccines, 12) virus testing framework, and 13) contact tracing). In addition, the introduced method learns to generate a graph to adjust the spatial dependences among different countries while forecasting the spread. We trained two models for short and long-term forecasts. <i>The first one</i> is trained to output one step in future with three previous timestamps of all features across the globe, whereas <i>the second model</i> is trained to output 10 steps in future. Overall, the trained models show high validation for forecasting the spread for each country for short and long-term forecasts, which makes the introduce method a useful tool to assist decision and policymaking for the different corners of the globe.</p>]]></description>
            <pubDate><![CDATA[2021-01-28T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Neural manifold under plasticity in a goal driven learning behaviour]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765840578391-8e0ccb38-28cc-4dd5-85a3-7125639f9571/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pcbi.1008621</link>
            <description><![CDATA[<p class="para" id="N65539">Neural activity is often low dimensional and dominated by only a few prominent neural covariation patterns. It has been hypothesised that these covariation patterns could form the building blocks used for fast and flexible motor control. Supporting this idea, recent experiments have shown that monkeys can learn to adapt their neural activity in motor cortex on a timescale of minutes, given that the change lies within the original low-dimensional subspace, also called neural manifold. However, the neural mechanism underlying this within-manifold adaptation remains unknown. Here, we show in a computational model that modification of recurrent weights, driven by a learned feedback signal, can account for the observed behavioural difference between within- and outside-manifold learning. Our findings give a new perspective, showing that recurrent weight changes do not necessarily lead to change in the neural manifold. On the contrary, successful learning is naturally constrained to a common subspace.</p><p class="para" id="N65542">It has been suggested that the coordinated activation of neurons might play an important role for movement execution. Whether such activity patterns are fixed or flexibly relearned remains matter of debate. It has been shown that monkeys can learn within minutes to adjust their neural activity, as long as they use the initial set of activity patterns. In contrast, monkeys needed several days and a sequential training procedure to learn completely new patterns. Here, we developed a computational model to investigate which biological features might lead to these experimental observations. Learning in our model is implemented through weight changes between neurons in a recurrently connected network. In order for these weight changes to improve the produced behaviour, an error signal is required which tells each neuron whether it should increase or decrease its activity in order to produce a movement closer to the target movement. We found that learning such an error signal is possible only in the first experimental condition, where monkeys needed to adapt their neural activity using already existing activity patterns. The learning of this error signal therefore poses a major constraint on what type of changes in neural activity can and can not be learned.</p>]]></description>
            <pubDate><![CDATA[2021-02-05T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Semi-Siamese U-Net for separation of lung and heart bioimpedance images: A simulation study of thorax EIT]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765839376205-26748928-3aa6-4452-a43d-f19fe1a71f71/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246071</link>
            <description><![CDATA[<p class="para" id="N65539">Electrical impedance tomography (EIT) is widely used for bedside monitoring of lung ventilation status. Its goal is to reflect the internal conductivity changes and estimate the electrical properties of the tissues in the thorax. However, poor spatial resolution affects EIT image reconstruction to the extent that the heart and lung-related impedance images are barely distinguishable. Several studies have attempted to tackle this problem, and approaches based on decomposition of EIT images using linear transformations have been developed, and recently, U-Net has become a prominent architecture for semantic segmentation. In this paper, we propose a novel semi-Siamese U-Net specifically tailored for EIT application. It is based on the state-of-the-art U-Net, whose structure is modified and extended, forming shared encoder with parallel decoders and has multi-task weighted losses added to adapt to the individual separation tasks. The trained semi-Siamese U-Net model was evaluated with a test dataset, and the results were compared with those of the classical U-Net in terms of Dice similarity coefficient and mean absolute error. Results showed that compared with the classical U-Net, semi-Siamese U-Net exhibited performance improvements of 11.37% and 3.2% in Dice similarity coefficient, and 3.16% and 5.54% in mean absolute error, in terms of heart and lung-impedance image separation, respectively.</p>]]></description>
            <pubDate><![CDATA[2021-02-02T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Performance improvement of machine learning techniques predicting the association of exacerbation of peak expiratory flow ratio with short term exposure level to indoor air quality using adult asthmatics clustered data]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765839208981-71def8b2-fe4d-47cb-bd0d-3a1e557ce28c/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244233</link>
            <description><![CDATA[<p class="para" id="N65539">Large-scale data sources, remote sensing technologies, and superior computing power have tremendously benefitted to environmental health study. Recently, various machine-learning algorithms were introduced to provide mechanistic insights about the heterogeneity of clustered data pertaining to the symptoms of each asthma patient and potential environmental risk factors. However, there is limited information on the performance of these machine learning tools. In this study, we compared the performance of ten machine-learning techniques. Using an advanced method of imbalanced sampling (IS), we improved the performance of nine conventional machine learning techniques predicting the association between exposure level to indoor air quality and change in patients’ peak expiratory flow rate (PEFR). We then proposed a deep learning method of transfer learning (TL) for further improvement in prediction accuracy. Our selected final prediction techniques (TL1_IS or TL2-IS) achieved a balanced accuracy median (interquartile range) of 66(56~76) % for TL1_IS and 68(63~78) % for TL2_IS. Precision levels for TL1_IS and TL2_IS were 68(62~72) % and 66(62~69) % while sensitivity levels were 58(50~67) % and 59(51~80) % from 25 patients which were approximately 1.08 (accuracy, precision) to 1.28 (sensitivity) times increased in terms of performance outcomes, compared to NN_IS. Our results indicate that the transfer machine learning technique with imbalanced sampling is a powerful tool to predict the change in PEFR due to exposure to indoor air including the concentration of particulate matter of 2.5 μm and carbon dioxide. This modeling technique is even applicable with small-sized or imbalanced dataset, which represents a personalized, real-world setting.</p>]]></description>
            <pubDate><![CDATA[2021-01-07T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Genome-wide prediction of topoisomerase II<i>β</i> binding by architectural factors and chromatin accessibility]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765838779416-084549e8-0b4f-4061-882c-4c7f1f5b0ff1/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pcbi.1007814</link>
            <description><![CDATA[<p class="para" id="N65539">DNA topoisomerase II-<i>β</i> (TOP2B) is fundamental to remove topological problems linked to DNA metabolism and 3D chromatin architecture, but its cut-and-reseal catalytic mechanism can accidentally cause DNA double-strand breaks (DSBs) that can seriously compromise genome integrity. Understanding the factors that determine the genome-wide distribution of TOP2B is therefore not only essential for a complete knowledge of genome dynamics and organization, but also for the implications of TOP2-induced DSBs in the origin of oncogenic translocations and other types of chromosomal rearrangements. Here, we conduct a machine-learning approach for the prediction of TOP2B binding using publicly available sequencing data. We achieve highly accurate predictions, with accessible chromatin and architectural factors being the most informative features. Strikingly, TOP2B is sufficiently explained by only three features: DNase I hypersensitivity, CTCF and cohesin binding, for which genome-wide data are widely available. Based on this, we develop a predictive model for TOP2B genome-wide binding that can be used across cell lines and species, and generate virtual probability tracks that accurately mirror experimental ChIP-seq data. Our results deepen our knowledge on how the accessibility and 3D organization of chromatin determine TOP2B function, and constitute a proof of principle regarding the <i>in silico</i> prediction of sequence-independent chromatin-binding factors.</p><p class="para" id="N65542">Type II DNA topoisomerases (TOP2) are a double-edged sword. They solve topological problems in the form of supercoiling, knots and tangles that inevitably accompany genome metabolism, but they do so at the cost of transiently cleaving DNA, with the risk that this entails for genome integrity, and the serious consequences for human health, such as neurodegeneration, developmental disorders or predisposition to cancer. A comprehensive analysis of TOP2 distribution throughout the genome is therefore essential for a deep understanding of its function and regulation, and how this can affect genome dynamics and stability. Here, we use machine learning to thoroughly explore genome-wide binding of TOP2B, a vertebrate TOP2 paralog that has been linked to genome organization and cancer-associated translocations. Our analysis shows that TOP2B-DNA binding can be accurately predicted exclusively using information on DNA accessibility and binding of genome-architecture factors. We show that such information is enough to generate virtual maps of TOP2B binding along the genome, which we validate with <i>de novo</i> experimental data. Our results highlight the importance of TOP2B for accessibility and 3D organization of chromatin, and show that computationally predicted TOP2 maps can be accurately obtained using minimal publicly available datasets, opening the door for their use in different organisms, cell types and conditions with experimental and/or clinical relevance.</p>]]></description>
            <pubDate><![CDATA[2021-01-19T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Optimal learning with excitatory and inhibitory synapses]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765838422474-325dd05d-78bf-43e9-a4ee-0f0455ba9f34/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pcbi.1008536</link>
            <description><![CDATA[<p class="para" id="N65539">Characterizing the relation between weight structure and input/output statistics is fundamental for understanding the computational capabilities of neural circuits. In this work, I study the problem of storing associations between analog signals in the presence of correlations, using methods from statistical mechanics. I characterize the typical learning performance in terms of the power spectrum of random input and output processes. I show that optimal synaptic weight configurations reach a capacity of 0.5 for any fraction of excitatory to inhibitory weights and have a peculiar synaptic distribution with a finite fraction of silent synapses. I further provide a link between typical learning performance and principal components analysis in single cases. These results may shed light on the synaptic profile of brain circuits, such as cerebellar structures, that are thought to engage in processing time-dependent signals and performing on-line prediction.</p><p class="para" id="N65542">A general analysis of learning with biological synaptic constraints in the presence of statistically structured signals is lacking. Here, analytical techniques from statistical mechanics are leveraged to analyze association storage between analog inputs and outputs with excitatory and inhibitory synaptic weights. The linear perceptron performance is characterized and a link is provided between the weight distribution and the correlations of input/output signals. This formalism can be used to predict the typical properties of perceptron solutions for single learning instances in terms of the principal component analysis of input and output data. This study provides a mean-field theory for sign-constrained regression of practical importance in neuroscience as well as in adaptive control applications.</p>]]></description>
            <pubDate><![CDATA[2020-12-28T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Identifying intentional injuries among children and adolescents based on Machine Learning]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765837850377-3ff52a85-6fd0-44bd-9b3e-f066c30bcb43/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245437</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Background</h3><p class="para" id="N65543">Compared to other studies, the injury monitoring of Chinese children and adolescents has captured a low level of intentional injuries on account of self-harm/suicide and violent attacks. Intentional injuries in children and adolescents have not been apparent from the data. It is possible that there has been a misclassification of existing intentional injuries, and there is a lack of research literature on the misclassification of intentional injuries. This study aimed to discuss the feasibility of discriminating the intention of injury based on Machine Learning (ML) modelling and provided ideas for understanding whether there was a misclassification of intentional injuries.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Methods</h3><p class="para" id="N65549">Information entropy was used to determine the correlation between variables and the intention of injury, and Naive Bayes (NB), Decision Tree (DT), Random Forest (RF), Adaboost algorithms and Deep Neural Networks (DNN) were used to create an intention of injury discrimination model. The models were compared by comprehensively testing the discrimination effect to determine stability and consistency.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Results</h3><p class="para" id="N65555">For the area under the ROC curve with different intentions of injuries, the NB model was 0.891, 0.880, and 0.897, respectively; the DT model was 0.870, 0.803, and 0.871, respectively; the RF model was 0.850, 0.809, and 0.845, respectively; the Adaboost model was 0.914, 0.846, and 0.914, respectively; the DNN model was 0.927, 0.835, and 0.934, respectively. In a comprehensive comparison of the five models, DNN and Adaboost models had higher values for the determination of the intention of injury. A discrimination of cases with unclear intentions of injury showed that on average, unintentional injuries, violent attacks, and self-harm/suicides accounted for 86.57%, 6.81%, and 6.62%, respectively.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Conclusion</h3><p class="para" id="N65561">It was feasible to use the ML algorithm to determine the injury intention of children and adolescents. The research suggested that the DNN and Adaboost models had higher values for the determination of the intention of injury. This study could build a foundation for transforming the model into a tool for rapid diagnosis and excavating potential intentional injuries of children and adolescents by widely collecting the influencing factors, extracting the influence variables characteristically, reducing the complexity and improving the performance of the models in the future.</p></div>]]></description>
            <pubDate><![CDATA[2021-01-20T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Machine learning for buildings’ characterization and power-law recovery of urban metrics]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765837559514-ec94f320-aad9-4d90-bf3c-5f1980526b31/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246096</link>
            <description><![CDATA[<p class="para" id="N65539">In this paper we focus on a critical component of the city: its building stock, which holds much of its socio-economic activities. In our case, the lack of a comprehensive database about their features and its limitation to a surveyed subset lead us to adopt data-driven techniques to extend our knowledge to the near-city-scale. Neural networks and random forests are applied to identify the buildings’ number of floors and construction periods’ dependencies on a set of shape features: area, perimeter, and height along with the annual electricity consumption, relying a surveyed data in the city of Beirut. The predicted results are then compared with established scaling laws of urban forms, which constitutes a further consistency check and validation of our workflow.</p>]]></description>
            <pubDate><![CDATA[2021-01-28T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[FOSTER—An R package for forest structure extrapolation]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765837400117-b55978c4-2114-411a-8ed7-70b99a67aa21/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244846</link>
            <description><![CDATA[<p class="para" id="N65539">The uptake of technologies such as airborne laser scanning (ALS) and more recently digital aerial photogrammetry (DAP) enable the characterization of 3-dimensional (3D) forest structure. These forest structural attributes are widely applied in the development of modern enhanced forest inventories. As an alternative to extensive ALS or DAP based forest inventories, regional forest attribute maps can be built from relationships between ALS or DAP and wall-to-wall satellite data products. To date, a number of different approaches exist, with varying code implementations using different programming environments and tailored to specific needs. With the motivation for open, simple and modern software, we present <tt>FOSTER</tt> (Forest Structure Extrapolation in R), a versatile and computationally efficient framework for modeling and imputation of 3D forest attributes. <tt>FOSTER</tt> derives spectral trends in remote sensing time series, implements a structurally guided sampling approach to sample these often spatially auto correlated datasets, to then allow a modelling approach (currently k-NN imputation) to extrapolate these 3D forest structure measures. The k-NN imputation approach that <tt>FOSTER</tt> implements has a number of benefits over conventional regression based approaches including lower bias and reduced over fitting. This paper provides an overview of the general framework followed by a demonstration of the performance and outputs of <tt>FOSTER</tt>. Two ALS-derived variables, the 95<sup>th</sup> percentile of first returns height (<i>elev_p95</i>) and canopy cover above mean height (<i>cover</i>), were imputed over a research forest in British Columbia, Canada with relative RMSE of 18.5% and 11.4% and relative bias of -0.6% and 1.4% respectively. The processing sequence developed within <tt>FOSTER</tt> represents an innovative and versatile framework that should be useful to researchers and managers alike looking to make forest management decisions over entire forest estates.</p>]]></description>
            <pubDate><![CDATA[2021-01-28T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Mapping climate discourse to climate opinion: An approach for augmenting surveys with social media to enhance understandings of climate opinion in the United States]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765837072349-733f8801-367d-4297-bd3b-595d7e4b2eb2/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245319</link>
            <description><![CDATA[<p class="para" id="N65539">Surveys are commonly used to quantify public opinions of climate change and to inform sustainability policies. However, conducting large-scale population-based surveys is often a difficult task due to time and resource constraints. This paper outlines a machine learning framework—grounded in statistical learning theory and natural language processing—to augment climate change opinion surveys with social media data. The proposed framework maps social media discourse to climate opinion surveys, allowing for discerning the regionally distinct topics and themes that contribute to climate opinions. The analysis reveals significant regional variation in the emergent social media topics associated with climate opinions. Furthermore, significant correlation is identified between social media discourse and climate attitude. However, the dependencies between topic discussion and climate opinion are not always intuitive and often require augmenting the analysis with a topic’s most frequent n-grams and most representative tweets to effectively interpret the relationship. Finally, the paper concludes with a discussion of how these results can be used in the policy framing process to quickly and effectively understand constituents’ opinions on critical issues.</p>]]></description>
            <pubDate><![CDATA[2021-01-14T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[A comparison of machine learning models versus clinical evaluation for mortality prediction in patients with sepsis]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765835621852-f0733dec-649d-447c-ab49-0488f979b5dd/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245157</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Introduction</h3><p class="para" id="N65543">Patients with sepsis who present to an emergency department (ED) have highly variable underlying disease severity, and can be categorized from low to high risk. Development of a risk stratification tool for these patients is important for appropriate triage and early treatment. The aim of this study was to develop machine learning models predicting 31-day mortality in patients presenting to the ED with sepsis and to compare these to internal medicine physicians and clinical risk scores.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Methods</h3><p class="para" id="N65549">A single-center, retrospective cohort study was conducted amongst 1,344 emergency department patients fulfilling sepsis criteria. Laboratory and clinical data that was available in the first two hours of presentation from these patients were randomly partitioned into a development (n = 1,244) and validation dataset (n = 100). Machine learning models were trained and evaluated on the development dataset and compared to internal medicine physicians and risk scores in the independent validation dataset. The primary outcome was 31-day mortality.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Results</h3><p class="para" id="N65555">A number of 1,344 patients were included of whom 174 (13.0%) died. Machine learning models trained with laboratory or a combination of laboratory + clinical data achieved an area-under-the ROC curve of 0.82 (95% CI: 0.80–0.84) and 0.84 (95% CI: 0.81–0.87) for predicting 31-day mortality, respectively. In the validation set, models outperformed internal medicine physicians and clinical risk scores in sensitivity (92% vs. 72% vs. 78%;p&lt;0.001,all comparisons) while retaining comparable specificity (78% vs. 74% vs. 72%;p&gt;0.02). The model had higher diagnostic accuracy with an area-under-the-ROC curve of 0.85 (95%CI: 0.78–0.92) compared to abbMEDS (0.63,0.54–0.73), mREMS (0.63,0.54–0.72) and internal medicine physicians (0.74,0.65–0.82).</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Conclusion</h3><p class="para" id="N65561">Machine learning models outperformed internal medicine physicians and clinical risk scores in predicting 31-day mortality. These models are a promising tool to aid in risk stratification of patients presenting to the ED with sepsis.</p></div>]]></description>
            <pubDate><![CDATA[2021-01-19T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Clinical outcome prediction from analysis of microelectrode recordings using deep learning in subthalamic deep brain stimulation for Parkinson`s disease]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765835165459-30828c4a-7abc-4dca-9f14-b73bfcf7b918/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244133</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Background</h3><p class="para" id="N65543">Deep brain stimulation (DBS) of the subthalamic nucleus (STN) is an effective treatment for improving the motor symptoms of advanced Parkinson’s disease (PD). Accurate positioning of the stimulation electrodes is necessary for better clinical outcomes.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Objective</h3><p class="para" id="N65549">We applied deep learning techniques to microelectrode recording (MER) signals to better predict motor function improvement, represented by the UPDRS part III scores, after bilateral STN DBS in patients with advanced PD. If we find the optimal stimulation point with MER by deep learning, we can improve the clinical outcome of STN DBS even under restrictions such as general anesthesia or non-cooperation of the patients.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Methods</h3><p class="para" id="N65555">In total, 696 4-second left-side MER segments from 34 patients with advanced PD who underwent bilateral STN DBS surgery under general anesthesia were included. We transformed the original signal into three wavelets of 1–50 Hz, 50–500 Hz, and 500–5,000 Hz. The wavelet-transformed MER was used for input data of the deep learning. The patients were divided into two groups, good response and moderate response groups, according to DBS on to off ratio of UPDRS part III score for the off-medication state, 6 months postoperatively. The ratio were used for output data in deep learning. The Visual Geometry Group (VGG)-16 model with a multitask learning algorithm was used to estimate the bilateral effect of DBS. Different ratios of the loss function in the task-specific layer were applied considering that DBS affects both sides differently.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Results</h3><p class="para" id="N65561">When we divided the MER signals according to the frequency, the maximal accuracy was higher in the 50–500 Hz group than in the 1–50 Hz and 500–5,000 Hz groups. In addition, when the multitask learning method was applied, the stability of the model was improved in comparison with single task learning. The maximal accuracy (80.21%) occurred when the right-to-left loss ratio was 5:1 or 6:1. The area under the curve (AUC) was 0.88 in the receiver operating characteristic (ROC) curve.</p></div><div class="section" id="sec005"><h3 class="BHead" id="nov000-5">Conclusion</h3><p class="para" id="N65567">Clinical improvements in PD patients who underwent bilateral STN DBS could be predicted based on a multitask deep learning-based MER analysis.</p></div>]]></description>
            <pubDate><![CDATA[2021-01-26T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Applying machine learning and geolocation techniques to social media data (Twitter) to develop a resource for urban planning]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765834922464-548ff4bd-0cd5-4098-80dd-702b6cf80af8/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244317</link>
            <description><![CDATA[<p class="para" id="N65539">With all the recent attention focused on big data, it is easy to overlook that basic vital statistics remain difficult to obtain in most of the world. What makes this frustrating is that private companies hold potentially useful data, but it is not accessible by the people who can use it to track poverty, reduce disease, or build urban infrastructure. This project set out to test whether we can transform an openly available dataset (Twitter) into a resource for urban planning and development. We test our hypothesis by creating road traffic crash location data, which is scarce in most resource-poor environments but essential for addressing the number one cause of mortality for children over five and young adults. The research project scraped 874,588 traffic related tweets in Nairobi, Kenya, applied a machine learning model to capture the occurrence of a crash, and developed an improved geoparsing algorithm to identify its location. We geolocate 32,991 crash reports in Twitter for 2012–2020 and cluster them into 22,872 unique crashes during this period. For a subset of crashes reported on Twitter, a motorcycle delivery service was dispatched in real-time to verify the crash and its location; the results show 92% accuracy. To our knowledge this is the first geolocated dataset of crashes for the city and allowed us to produce the first crash map for Nairobi. Using a spatial clustering algorithm, we are able to locate portions of the road network (&lt;1%) where 50% of the crashes identified occurred. Even with limitations in the representativeness of the data, the results can provide urban planners with useful information that can be used to target road safety improvements where resources are limited. The work shows how twitter data might be used to create other types of essential data for urban planning in resource poor environments.</p>]]></description>
            <pubDate><![CDATA[2021-02-03T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[SafeNET: Initial development and validation of a real-time tool for predicting mortality risk at the time of hospital transfer to a higher level of care]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765834893282-fcde5bf8-d052-4b09-a03f-0cc196a44828/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246669</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Background</h3><p class="para" id="N65543">Processes for transferring patients to higher acuity facilities lack a standardized approach to prognostication, increasing the risk for low value care that imposes significant burdens on patients and their families with unclear benefits. We sought to develop a rapid and feasible tool for predicting mortality using variables readily available at the time of hospital transfer.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Methods and findings</h3><p class="para" id="N65549">All work was carried out at a single, large, multi-hospital integrated healthcare system. We used a retrospective cohort for model development consisting of patients aged 18 years or older transferred into the healthcare system from another hospital, hospice, skilled nursing or other healthcare facility with an admission priority of direct emergency admit. The cohort was randomly divided into training and test sets to develop first a 54-variable, and then a 14-variable gradient boosting model to predict the primary outcome of all cause in-hospital mortality. Secondary outcomes included 30-day and 90-day mortality and transition to comfort measures only or hospice care. For model validation, we used a prospective cohort consisting of all patients transferred to a single, tertiary care hospital from one of the 3 referring hospitals, excluding patients transferred for myocardial infarction or maternal labor and delivery. Prospective validation was performed by using a web-based tool to calculate the risk of mortality at the time of transfer. Observed outcomes were compared to predicted outcomes to assess model performance.</p><p class="para" id="N65551">The development cohort included 20,985 patients with 1,937 (9.2%) in-hospital mortalities, 2,884 (13.7%) 30-day mortalities, and 3,899 (18.6%) 90-day mortalities. The 14-variable gradient boosting model effectively predicted in-hospital, 30-day and 90-day mortality (c = 0.903 [95% CI:0.891–0.916]), c = 0.877 [95% CI:0.864–0.890]), and c = 0.869 [95% CI:0.857–0.881], respectively). The tool was proven feasible and valid for bedside implementation in a prospective cohort of 679 sequentially transferred patients for whom the bedside nurse calculated a SafeNET score at the time of transfer, taking only 4–5 minutes per patient with discrimination consistent with the development sample for in-hospital, 30-day and 90-day mortality (c = 0.836 [95%CI: 0.751–0.921], 0.815 [95% CI: 0.730–0.900], and 0.794 [95% CI: 0.725–0.864], respectively).</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Conclusions</h3><p class="para" id="N65557">The SafeNET algorithm is feasible and valid for real-time, bedside mortality risk prediction at the time of hospital transfer. Work is ongoing to build pathways triggered by this score that direct needed resources to the patients at greatest risk of poor outcomes.</p></div>]]></description>
            <pubDate><![CDATA[2021-02-08T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Supervised machine learning for automated classification of human Wharton’s Jelly cells and mechanosensory hair cells]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765834668343-0da76195-fcdb-4b25-9b99-c720ac47e57e/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245234</link>
            <description><![CDATA[<p class="para" id="N65539">Tissue engineering and gene therapy strategies offer new ways to repair permanent damage to mechanosensory hair cells (MHCs) by differentiating human Wharton’s Jelly cells (HWJCs). Conventionally, these strategies require the classification of each cell as differentiated or undifferentiated. Automated classification tools, however, may serve as a novel method to rapidly classify these cells. In this paper, images from previous work, where HWJCs were differentiated into MHC-like cells, were examined. Various cell features were extracted from these images, and those which were pertinent to classification were identified. Different machine learning models were then developed, some using all extracted data and some using only certain features. To evaluate model performance, the area under the curve (AUC) of the receiver operating characteristic curve was primarily used. This paper found that limiting algorithms to certain features consistently improved performance. The top performing model, a voting classifier model consisting of two logistic regressions, a support vector machine, and a random forest classifier, obtained an AUC of 0.9638. Ultimately, this paper illustrates the viability of a novel machine learning pipeline to automate the classification of undifferentiated and differentiated cells. In the future, this research could aid in automated strategies that determine the viability of MHC-like cells after differentiation.</p>]]></description>
            <pubDate><![CDATA[2021-01-08T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Bone strain index as a predictor of further vertebral fracture in osteoporotic women: An artificial intelligence-based analysis]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765833679092-ae2d1f38-6246-453d-89f4-8fdf7c8ad1d5/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245967</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Background</h3><p class="para" id="N65543">Osteoporosis is an asymptomatic disease of high prevalence and incidence, leading to bone fractures burdened by high mortality and disability, mainly when several subsequent fractures occur. A fragility fracture predictive model, Artificial Intelligence-based, to identify dual X-ray absorptiometry (DXA) variables able to characterise those patients who are prone to further fractures called Bone Strain Index, was evaluated in this study.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Methods</h3><p class="para" id="N65549">In a prospective, longitudinal, multicentric study 172 female outpatients with at least one vertebral fracture at the first observation were enrolled. They performed a spine X-ray to calculate spine deformity index (SDI) and a lumbar and femoral DXA scan to assess bone mineral density (BMD) and bone strain index (BSI) at baseline and after a follow-up period of 3 years in average. At the end of the follow-up, 93 women developed a further vertebral fracture. The further vertebral fracture was considered as one unit increase of SDI. We assessed the predictive capacity of supervised Artificial Neural Networks (ANNs) to distinguish women who developed a further fracture from those without it, and to detect those variables providing the maximal amount of relevant information to discriminate the two groups. ANNs choose appropriate input data automatically (TWIST-system, Training With Input Selection and Testing). Moreover, we built a semantic connectivity map usingthe Auto Contractive Map to provide further insights about the convoluted connections between the osteoporotic variables under consideration and the two scenarios (further fracture vs no further fracture).</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Results</h3><p class="para" id="N65555">TWIST system selected 5 out of 13 available variables: age, menopause age, BMI, FTot BMC, FTot BSI. With training testing procedure, ANNs reached predictive accuracy of 79.36%, with a sensitivity of 75% and a specificity of 83.72%.</p><p class="para" id="N65557">The semantic connectivity map highlighted the role of BSI in predicting the risk of a further fracture.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Conclusions</h3><p class="para" id="N65563">Artificial Intelligence is a useful method to analyse a complex system like that regarding osteoporosis, able to identify patients prone to a further fragility fracture. BSI appears to be a useful DXA index in identifying those patients who are at risk of further vertebral fractures.</p></div>]]></description>
            <pubDate><![CDATA[2021-02-08T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Predicting mortality of patients with acute kidney injury in the ICU using XGBoost model]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765833135227-f8e4f280-1eb0-4821-8a4f-84eac047fbe2/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246306</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Purpose</h3><p class="para" id="N65543">The goal of this study <b>is</b> to construct a mortality prediction model using the XGBoot (eXtreme Gradient Boosting) decision tree model for AKI (acute kidney injury) patients in the ICU (intensive care unit), and to compare its performance with that of three other machine learning models.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Methods</h3><p class="para" id="N65552">We used the eICU Collaborative Research Database (eICU-CRD) for model development and performance comparison. The prediction performance of the XGBoot model was compared with the other three machine learning models. These models included LR (logistic regression), SVM (support vector machines), and RF (random forest). In the model comparison, the AUROC (area under receiver operating curve), accuracy, precision, recall, and F1 score were used to evaluate the predictive performance of each model.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Results</h3><p class="para" id="N65558">A total of 7548 AKI patients were analyzed in this study. The overall in-hospital mortality of AKI patients was 16.35%. The best performing algorithm in this study was XGBoost with the highest AUROC (0.796, p &lt; 0.01), F1(0.922, p &lt; 0.01) and accuracy (0.860). The precision (0.860) and recall (0.994) of the XGBoost model rank second among the four models.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Conclusion</h3><p class="para" id="N65564">XGBoot model had obvious advantages of performance compared to the other machine learning models. This will be helpful for risk identification and early intervention for AKI patients at risk of death.</p></div>]]></description>
            <pubDate><![CDATA[2021-02-04T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Is it feasible to detect FLOSS version release events from textual messages? A case study on Stack Overflow]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765833006798-59188049-f829-4b58-a3ae-23a260d2e033/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246464</link>
            <description><![CDATA[<p class="para" id="N65539">Topic Detection and Tracking (TDT) is a very active research question within the area of text mining, generally applied to news feeds and Twitter datasets, where topics and events are detected. The notion of “event” is broad, but typically it applies to occurrences that can be detected from a single post or a message. Little attention has been drawn to what we call “micro-events”, which, due to their nature, cannot be detected from a single piece of textual information. The study investigates the feasibility of micro-event detection on textual data using a sample of messages from the Stack Overflow Q&amp;A platform and Free/Libre Open Source Software (FLOSS) version releases from Libraries.io dataset. We build pipelines for detection of micro-events using three different estimators whose parameters are optimized using a grid search approach. We consider two feature spaces: LDA topic modeling with sentiment analysis, and hSBM topics with sentiment analysis. The feature spaces are optimized using the recursive feature elimination with cross validation (RFECV) strategy. In our experiments we investigate whether there is a characteristic change in the topics distribution or sentiment features before or after micro-events take place and we thoroughly evaluate the capacity of each variant of our analysis pipeline to detect micro-events. Additionally, we perform a detailed statistical analysis of the models, including influential cases, variance inflation factors, validation of the linearity assumption, pseudo <i>R</i><sup>2</sup> measures and no-information rate. Finally, in order to study limits of micro-event detection, we design a method for generating micro-event synthetic datasets with similar properties to the real-world data, and use them to identify the micro-event detectability threshold for each of the evaluated classifiers.</p>]]></description>
            <pubDate><![CDATA[2021-02-04T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[mbkmeans: Fast clustering for single cell data using mini-batch <i>k</i>-means]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765824662564-0cd77c6e-42f6-4efe-b606-39212ffac8f5/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pcbi.1008625</link>
            <description><![CDATA[<p class="para" id="N65539">Single-cell RNA-Sequencing (scRNA-seq) is the most widely used high-throughput technology to measure genome-wide gene expression at the single-cell level. One of the most common analyses of scRNA-seq data detects distinct subpopulations of cells through the use of unsupervised clustering algorithms. However, recent advances in scRNA-seq technologies result in current datasets ranging from thousands to millions of cells. Popular clustering algorithms, such as <i>k</i>-means, typically require the data to be loaded entirely into memory and therefore can be slow or impossible to run with large datasets. To address this problem, we developed the <i>mbkmeans</i> R/Bioconductor package, an open-source implementation of the mini-batch <i>k</i>-means algorithm. Our package allows for on-disk data representations, such as the common HDF5 file format widely used for single-cell data, that do not require all the data to be loaded into memory at one time. We demonstrate the performance of the <i>mbkmeans</i> package using large datasets, including one with 1.3 million cells. We also highlight and compare the computing performance of <i>mbkmeans</i> against the standard implementation of <i>k</i>-means and other popular single-cell clustering methods. Our software package is available in Bioconductor at https://bioconductor.org/packages/mbkmeans.</p><p class="para" id="N65542">We developed the <i>mbkmeans</i> package (https://bioconductor.org/packages/mbkmeans) in Bioconductor, an open-source implementation of the mini-batch <i>k</i>-means algorithm. Our package allows for on-disk data representations, such as the common HDF5 file format widely used for single-cell data, that do not require all the data to be loaded into memory at one time.</p>]]></description>
            <pubDate><![CDATA[2021-01-26T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Evaluating Deep Learning models for predicting ALK-5 inhibition]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765824566206-3c0dfd64-61cd-4d5d-b7fc-6ec59c3907d3/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246126</link>
            <description><![CDATA[<p class="para" id="N65539">Computational methods have been widely used in drug design. The recent developments in machine learning techniques and the ever-growing chemical and biological databases are fertile ground for discoveries in this area. In this study, we evaluated the performance of Deep Learning models in comparison to Random Forest, and Support Vector Regression for predicting the biological activity (pIC<sub>50</sub>) of ALK-5 inhibitors as candidates to treat cancer. The generalization power of the models was assessed by internal and external validation procedures. A deep neural network model obtained the best performance in this comparative study, achieving a coefficient of determination of 0.658 on the external validation set with mean square error and mean absolute error of 0.373 and 0.450, respectively. Additionally, the relevance of the chemical descriptors for the prediction of biological activity was estimated using Permutation Importance. We can conclude that the forecast model obtained by the deep neural network is suitable for the problem and can be employed to predict the biological activity of new ALK-5 inhibitors.</p>]]></description>
            <pubDate><![CDATA[2021-01-28T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Recognition of industrial machine parts based on transfer learning with convolutional neural network]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765824214591-9faabc6a-0db7-4e9b-b384-d57b1d851126/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245735</link>
            <description><![CDATA[<p class="para" id="N65539">As the industry gradually enters the stage of unmanned and intelligent, factories in the future need to realize intelligent monitoring and diagnosis and maintenance of parts and components. In order to achieve this goal, it is first necessary to accurately identify and classify the parts in the factory. However, the existing literature rarely studies the classification and identification of parts of the entire factory. Due to the lack of existing data samples, this paper studies the identification and classification of small samples of industrial machine parts. In order to solve this problem, this paper establishes a convolutional neural network model based on the InceptionNet-V3 pretrained model through migration learning. Through experimental design, the influence of data expansion, learning rate and optimizer algorithm on the model effectiveness is studied, and the optimal model was finally determined, and the test accuracy rate reaches 99.74%. By comparing with the accuracy of other classifiers, the experimental results prove that the convolutional neural network model based on transfer learning can effectively solve the problem of recognition and classification of industrial machine parts with small samples and the idea of transfer learning can also be further promoted.</p>]]></description>
            <pubDate><![CDATA[2021-01-28T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Accurate prediction of clinical stroke scales and improved biomarkers of motor impairment from robotic measurements]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765821823405-1a732d64-290e-49e9-836f-c51db9153780/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245874</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Objective</h3><p class="para" id="N65543">One of the greatest challenges in clinical trial design is dealing with the subjectivity and variability introduced by human raters when measuring clinical end-points. We hypothesized that robotic measures that capture the kinematics of human movements collected longitudinally in patients after stroke would bear a significant relationship to the ordinal clinical scales and potentially lead to the development of more sensitive motor biomarkers that could improve the efficiency and cost of clinical trials.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Materials and methods</h3><p class="para" id="N65549">We used clinical scales and a robotic assay to measure arm movement in 208 patients 7, 14, 21, 30 and 90 days after acute ischemic stroke at two separate clinical sites. The robots are low impedance and low friction interactive devices that precisely measure speed, position and force, so that even a hemiparetic patient can generate a complete measurement profile. These profiles were used to develop predictive models of the clinical assessments employing a combination of artificial ant colonies and neural network ensembles.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Results</h3><p class="para" id="N65555">The resulting models replicated commonly used clinical scales to a cross-validated R<sup>2</sup> of 0.73, 0.75, 0.63 and 0.60 for the Fugl-Meyer, Motor Power, NIH stroke and modified Rankin scales, respectively. Moreover, when suitably scaled and combined, the robotic measures demonstrated a significant increase in effect size from day 7 to 90 over historical data (1.47 versus 0.67).</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Discussion and conclusion</h3><p class="para" id="N65564">These results suggest that it is possible to derive surrogate biomarkers that can significantly reduce the sample size required to power future stroke clinical trials.</p></div>]]></description>
            <pubDate><![CDATA[2021-01-29T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Application of artificial neural network and support vector regression in predicting mass of ber fruits (<i>Ziziphus mauritiana</i> Lamk.) based on fruit axial dimensions]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765820995219-8d7a4144-7445-495a-ba9b-1f8c99f85ae1/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245228</link>
            <description><![CDATA[<p class="para" id="N65539">Fruit quality attributes are important factors for designing a market for agricultural goods and commodities. Support vector regression (SVR), MLR, and ANN models were established to predict the mass of ber fruits (Ziziphus mauritiana Lamk.) based on the axial dimensions of the fruit from manual measurements of fruit length, minor fruit diameter, and maximum fruit diameter of four ber cultivars. The precision and accuracy of the established models were assessed given their predicted values. The results revealed that using the validation dataset, the developed ANN (R<sup>2</sup> = 0.9771; root mean square error [RMSE] = 1.8479 g) and SVR (R<sup>2</sup> = 0.9947; RMSE = 1.8814 g) models produced better results when predicting ber fruit mass than those obtained by the MLR model (R<sup>2</sup> = 0.4614; RMSE = 11.3742 g). In estimating ber fruit mass, the established SVR and ANN models produced more precise prediction values than those produced by the MLR model; however, the performance differences between the SVR and ANN models were not clear.</p>]]></description>
            <pubDate><![CDATA[2021-01-07T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Forecasting hand-foot-and-mouth disease cases using wavelet-based SARIMA–NNAR hybrid model]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765819297614-d5dcc575-59e8-41f2-82ca-e7fa1e6e0694/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0246673</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Background</h3><p class="para" id="N65543">Hand-foot-and-mouth disease_(HFMD) is one of the most typical diseases in children that is associated with high morbidity. Reliable forecasting is crucial for prevention and control. Recently, hybrid models have become popular, and wavelet analysis has been widely performed. Better prediction accuracy may be achieved using wavelet-based hybrid models. Thus, our aim is to forecast number of HFMD cases with wavelet-based hybrid models.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Materials and methods</h3><p class="para" id="N65549">We fitted a wavelet-based seasonal autoregressive integrated moving average (SARIMA)–neural network nonlinear autoregressive (NNAR) hybrid model with HFMD weekly cases from 2009 to 2016 in Zhengzhou, China. Additionally, a single SARIMA model, simplex NNAR model, and pure SARIMA–NNAR hybrid model were established for comparison and estimation.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Results</h3><p class="para" id="N65555">The wavelet-based SARIMA–NNAR hybrid model demonstrates excellent performance whether in fitting or forecasting compared with other models. Its fitted and forecasting time series are similar to the actual observed time series.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Conclusions</h3><p class="para" id="N65561">The wavelet-based SARIMA–NNAR hybrid model fitted in this study is suitable for forecasting the number of HFMD cases. Hence, it will facilitate the prevention and control of HFMD.</p></div>]]></description>
            <pubDate><![CDATA[2021-02-05T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[On transformative adaptive activation functions in neural networks for gene expression inference]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765819026056-96b8f1a1-8455-4dd7-a3e5-86bacd79607c/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0243915</link>
            <description><![CDATA[<p class="para" id="N65539">Gene expression profiling was made more cost-effective by the NIH LINCS program that profiles only ∼1, 000 selected landmark genes and uses them to reconstruct the whole profile. The D–GEX method employs neural networks to infer the entire profile. However, the original D–GEX can be significantly improved. We propose a novel transformative adaptive activation function that improves the gene expression inference even further and which generalizes several existing adaptive activation functions. Our improved neural network achieves an average mean absolute error of 0.1340, which is a significant improvement over our reimplementation of the original D–GEX, which achieves an average mean absolute error of 0.1637. The proposed transformative adaptive function enables a significantly more accurate reconstruction of the full gene expression profiles with only a small increase in the complexity of the model and its training procedure compared to other methods.</p>]]></description>
            <pubDate><![CDATA[2021-01-14T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[The geometric approach to human stress based on stress-related surrogate measures]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765799870810-4d9b9bb2-2e51-4878-a90f-94fc935d50ce/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0219414</link>
            <description><![CDATA[<p class="para" id="N65539">We present a <i>predictive Geometric Stress Index</i> (pGSI) and its relation to behavioural Entropy (bE<div class="imageVideo"><img src="" alt=""/></div>). bE<div class="imageVideo"><img src="" alt=""/></div> is a measure of the complexity of an organism’s reactivity to stressors yielding patterns based on different behavioural and physiological variables selected as Surrogate Markers of Stress (SMS). We present a relationship between pGSI and bE<div class="imageVideo"><img src="" alt=""/></div> in terms of a power law model. This nonlinear relationship describes congruences in complexity derived from analyses of observable and measurable SMS based patterns interpreted as stress. The adjective geometric refers to subdivision(s) of the domain derived from two SMS (heart rate variability and steps frequency) with respect to a positive/negative binary perceptron based on a third SMS (blood oxygenation). The presented power law allows for both quantitative and qualitative evaluations of the consequences of stress measured by pGSI. In particular, we show that elevated stress levels in terms of pGSI leads to a decrease of the bE<div class="imageVideo"><img src="" alt=""/></div> of the blood oxygenation, measured by peripheral blood oxygenation S<sub><i>p</i></sub>O<sub>2</sub> as a model of SMS.</p>]]></description>
            <pubDate><![CDATA[2021-01-25T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Machine learning and statistics to qualify environments through multi-traits in <i>Coffea arabica</i>]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765793682629-ee5c6b6f-8a9f-4556-85b6-da6aa8b21437/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245298</link>
            <description><![CDATA[<p class="para" id="N65539">Several factors such as genotype, environment, and post-harvest processing can affect the responses of important traits in the coffee production chain. Determining the influence of these factors is of great relevance, as they can be indicators of the characteristics of the coffee produced. The most efficient models choice to be applied should take into account the variety of information and the particularities of each biological material. This study was developed to evaluate statistical and machine learning models that would better discriminate environments through multi-traits of coffee genotypes and identify the main agronomic and beverage quality traits responsible for the variation of the environments. For that, 31 morpho-agronomic and post-harvest traits were evaluated, from field experiments installed in three municipalities in the Matas de Minas region, in the State of Minas Gerais, Brazil. Two types of post-harvest processing were evaluated: natural and pulped. The apparent error rate was estimated for each method. The Multilayer Perceptron and Radial Basis Function networks were able to discriminate the coffee samples in multi-environment more efficiently than the other methods, identifying differences in multi-traits responses according to the production sites and type of post-harvest processing. The local factors did not present specific traits that favored the severity of diseases and differentiated vegetative vigor. Sensory traits acidity and fragrance/aroma score also made little contribution to the discrimination process, indicating that acidity and fragrance/aroma are characteristic of coffee produced and all coffee samples evaluated are of the special type in the Mata of Minas region. The main traits responsible for the differentiation of production sites are plant height, fruit size, and bean production. The sensory trait "Body" is the main one to discriminate the form of post-harvest processing.</p>]]></description>
            <pubDate><![CDATA[2021-01-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[COVID-19: Short-term forecast of ICU beds in times of crisis]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765792970521-b27372a7-a38f-41b5-9e29-4ed2af8ce68f/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245272</link>
            <description><![CDATA[<p class="para" id="N65539">By early May 2020, the number of new COVID-19 infections started to increase rapidly in Chile, threatening the ability of health services to accommodate all incoming cases. Suddenly, ICU capacity planning became a first-order concern, and the health authorities were in urgent need of tools to estimate the demand for urgent care associated with the pandemic. In this article, we describe the approach we followed to provide such demand forecasts, and we show how the use of analytics can provide relevant support for decision making, even with incomplete data and without enough time to fully explore the numerical properties of all available forecasting methods. The solution combines autoregressive, machine learning and epidemiological models to provide a short-term forecast of ICU utilization at the regional level. These forecasts were made publicly available and were actively used to support capacity planning. Our predictions achieved average forecasting errors of 4% and 9% for one- and two-week horizons, respectively, outperforming several other competing forecasting models.</p>]]></description>
            <pubDate><![CDATA[2021-01-13T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Behavioral structure of users in cryptocurrency market]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765789675120-9622bb89-0289-4e5a-ad11-0bde7182d1e9/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0242600</link>
            <description><![CDATA[<p class="para" id="N65539">Human behavior as they engaged in financial activities is intimately connected to the observed market dynamics. Despite many existing theories and studies on the fundamental motivations of the behavior of humans in financial systems, there is still limited empirical deduction of the behavioral compositions of the financial agents from a detailed market analysis. Blockchain technology has provided an avenue for the latter investigation with its voluminous data and its transparency of financial transactions. It has enabled us to perform empirical inference on the behavioral patterns of users in the market, which we explore in the bitcoin and ethereum cryptocurrency markets. In our study, we first determine various properties of the bitcoin and ethereum users by a temporal complex network analysis. After which, we develop methodology by combining <i>k</i>-means clustering and Support Vector Machines to derive behavioral types of users in the two cryptocurrency markets. Interestingly, we found four distinct strategies that are common in both markets: optimists, pessimists, positive traders and negative traders. The composition of user behavior is remarkably different between the bitcoin and ethereum market during periods of local price fluctuations and large systemic events. We observe that bitcoin (ethereum) users tend to take a short-term (long-term) view of the market during the local events. For the large systemic events, ethereum (bitcoin) users are found to consistently display a greater sense of pessimism (optimism) towards the future of the market.</p>]]></description>
            <pubDate><![CDATA[2021-01-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Implementing a high-efficiency similarity analysis approach for firmware code]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765789639419-5a6dd641-ea3f-4dea-8294-ea9f3e4bdab8/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245098</link>
            <description><![CDATA[<p class="para" id="N65539">The rapid expansion of the open-source community has shortened the software development cycle, but the spread of vulnerabilities has been accelerated, especially in the field of the Internet of Things. In recent years, the frequency of attacks against connected devices is increasing exponentially; thus, the vulnerabilities are more serious in nature. The state-of-the-art firmware security inspection technologies, such as methods based on machine learning and graph theory, find similar applications depending on the known vulnerabilities but cannot do anything without detailed information about the vulnerabilities. Moreover, model training, which is necessary for the machine learning technologies, requires a significant amount of time and data, resulting in low efficiency and poor extensibility. Aiming at the above shortcomings, a high-efficiency similarity analysis approach for firmware code is proposed in this study. First, the function control flow features and data flow features are extracted from the functions of the firmware and of the vulnerabilities, and the features are used to calculate the SimHash of the functions. The mass storage and fast query capabilities of the SimHash are implemented by the pigeonhole principle. Second, the similarity function pairs are analyzed in detail within and among the basic blocks. Within the basic blocks, the symbolic execution is used to generate the basic block semantic information, and the constraint solver is used to determine the semantic equivalence. Among the basic blocks, the local control flow graphs are analyzed to obtain their similarity. Then, we implemented a prototype and present the evaluation. The evaluation results demonstrate that the proposed approach can implement large-scale firmware function similarity analysis. It can also get the location of the real-world firmware patch without vulnerability function information. Finally, we compare our method with existing methods. The comparison results demonstrate that our method is more efficient and accurate than the Gemini and StagedMethod. More than 90% of the firmware functions can be indexed within 0.1 s, while the search time of 100,000 firmware functions is less than 2 s.</p>]]></description>
            <pubDate><![CDATA[2021-01-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Efficient gene expression signature for a breast cancer immuno-subtype]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765789035201-a2d21c3d-0c09-4af8-be46-fccb427f6b9f/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245215</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Motivation and background</h3><p class="para" id="N65543">The patient’s immune system plays an important role in cancer pathogenesis, prognosis and susceptibility to treatment. Recent work introduced an immune related breast cancer. This subtyping is based on the expression profiles of the tumor samples. Specifically, one study showed that analyzing 658 genes can lead to a signature for subtyping tumors. Furthermore, this classification is independent of other known molecular and clinical breast cancer subtyping. Finally, that study shows that the suggested subtyping has significant prognostic implications.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Results</h3><p class="para" id="N65549">In this work we develop an efficient signature associated with survival in breast cancer. We begin by developing a more efficient signature for the above-mentioned breast cancer immune-based subtyping. This signature represents better performance with a set of 579 genes that obtains an improved Area Under Curve (AUC). We then determine a set of 193 genes and an associated classification rule that yield subtypes with a much stronger statistically significant (log rank p-value &lt; 2 × 10<sup>−4</sup> in a test cohort) difference in survival. To obtain these improved results we develop a feature selection process that matches the high-dimensionality character of the data and the dual performance objectives, driven by survival and anchored by the literature subtyping.</p></div>]]></description>
            <pubDate><![CDATA[2021-01-12T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Robust radiogenomics approach to the identification of <i>EGFR</i> mutations among patients with NSCLC from three different countries using topologically invariant Betti numbers]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765788652680-5ce53f1a-278f-4020-b257-42d5986c73a5/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244354</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Objectives</h3><p class="para" id="N65543">To propose a novel robust radiogenomics approach to the identification of epidermal growth factor receptor (<i>EGFR</i>) mutations among patients with non-small cell lung cancer (NSCLC) using Betti numbers (BNs).</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Materials and methods</h3><p class="para" id="N65552">Contrast enhanced computed tomography (CT) images of 194 multi-racial NSCLC patients (79 <i>EGFR</i> mutants and 115 wildtypes) were collected from three different countries using 5 manufacturers’ scanners with a variety of scanning parameters. Ninety-nine cases obtained from the University of Malaya Medical Centre (UMMC) in Malaysia were used for training and validation procedures. Forty-one cases collected from the Kyushu University Hospital (KUH) in Japan and fifty-four cases obtained from The Cancer Imaging Archive (TCIA) in America were used for a test procedure. Radiomic features were obtained from BN maps, which represent topologically invariant heterogeneous characteristics of lung cancer on CT images, by applying histogram- and texture-based feature computations. A BN-based signature was determined using support vector machine (SVM) models with the best combination of features that maximized a robustness index (RI) which defined a higher total area under receiver operating characteristics curves (AUCs) and lower difference of AUCs between the training and the validation. The SVM model was built using the signature and optimized in a five-fold cross validation. The BN-based model was compared to conventional original image (OI)- and wavelet-decomposition (WD)-based models with respect to the RI between the validation and the test.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Results</h3><p class="para" id="N65561">The BN-based model showed a higher RI of 1.51 compared with the models based on the OI (RI: 1.33) and the WD (RI: 1.29).</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Conclusion</h3><p class="para" id="N65567">The proposed model showed higher robustness than the conventional models in the identification of <i>EGFR</i> mutations among NSCLC patients. The results suggested the robustness of the BN-based approach against variations in image scanner/scanning parameters.</p></div>]]></description>
            <pubDate><![CDATA[2021-01-11T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Insights into mobile health application market via a content analysis of marketplace data with machine learning]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765768032727-f3b0cd7f-d6ed-4ac2-a58f-ec2a1ed30491/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244302</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Background</h3><p class="para" id="N65543">Despite the benefits offered by an abundance of health applications promoted on app marketplaces (e.g., Google Play Store), the wide adoption of mobile health and e-health apps is yet to come.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Objective</h3><p class="para" id="N65549">This study aims to investigate the current landscape of smartphone apps that focus on improving and sustaining health and wellbeing. Understanding the categories that popular apps focus on and the relevant features provided to users, which lead to higher user scores and downloads will offer insights to enable higher adoption in the general populace. This study on 1,000 mobile health applications aims to shed light on the reasons why particular apps are liked and adopted while many are not.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Methods</h3><p class="para" id="N65555">User-generated data (i.e. review scores) and company-generated data (i.e. app descriptions) were collected from app marketplaces and manually coded and categorized by two researchers. For analysis, Artificial Neural Networks, Random Forest and Naïve Bayes Artificial Intelligence algorithms were used.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Results</h3><p class="para" id="N65561">The analysis led to features that attracted more download behavior and higher user scores. The findings suggest that apps that mention a privacy policy or provide videos in description lead to higher user scores, whereas free apps with in-app purchase possibilities, social networking and sharing features and feedback mechanisms lead to higher number of downloads. Moreover, differences in user scores and the total number of downloads are detected in distinct subcategories of mobile health apps.</p></div><div class="section" id="sec005"><h3 class="BHead" id="nov000-5">Conclusion</h3><p class="para" id="N65567">This study contributes to the current knowledge of m-health application use by reviewing mobile health applications using content analysis and machine learning algorithms. The content analysis adds significant value by providing classification, keywords and factors that influence download behavior and user scores in a m-health context.</p></div>]]></description>
            <pubDate><![CDATA[2021-01-06T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Reverse annealing for nonnegative/binary matrix factorization]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765767584928-9a6469f8-821e-4540-a568-413f48232fe7/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244026</link>
            <description><![CDATA[<p class="para" id="N65539">It was recently shown that quantum annealing can be used as an effective, fast subroutine in certain types of matrix factorization algorithms. The quantum annealing algorithm performed best for quick, approximate answers, but performance rapidly plateaued. In this paper, we utilize reverse annealing instead of forward annealing in the quantum annealing subroutine for nonnegative/binary matrix factorization problems. After an initial global search with forward annealing, reverse annealing performs a series of local searches that refine existing solutions. The combination of forward and reverse annealing significantly improves performance compared to forward annealing alone for all but the shortest run times.</p>]]></description>
            <pubDate><![CDATA[2021-01-06T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Recurrent disease progression networks for modelling risk trajectory of heart failure]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765767063164-db25a945-16fd-4d16-b649-02ba648213f2/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0245177</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Motivation</h3><p class="para" id="N65543">Recurrent neural networks (RNN) are powerful frameworks to model medical time series records. Recent studies showed improved accuracy of predicting future medical events (e.g., readmission, mortality) by leveraging large amount of high-dimensional data. However, very few studies have explored the ability of RNN in predicting long-term trajectories of recurrent events, which is more informative than predicting one single event in directing medical intervention.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Methods</h3><p class="para" id="N65549">In this study, we focus on heart failure (HF) which is the leading cause of death among cardiovascular diseases. We present a novel RNN framework named Deep Heart-failure Trajectory Model (DHTM) for modelling the long-term trajectories of recurrent HF. DHTM auto-regressively predicts the future HF onsets of each patient and uses the predicted HF as input to predict the HF event at the next time point. Furthermore, we propose an augmented DHTM named DHTM+C (where “C” stands for co-morbidities), which jointly predicts both the HF and a set of acute co-morbidities diagnoses. To efficiently train the DHTM+C model, we devised a novel RNN architecture to model disease progression implicated in the co-morbidities.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Results</h3><p class="para" id="N65555">Our deep learning models confers higher prediction accuracy for both the next-step HF prediction and the HF trajectory prediction compared to the baseline non-neural network models and the baseline RNN model. Compared to DHTM, DHTM+C is able to output higher probability of HF for high-risk patients, even in cases where it is only given less than 2 years of data to predict over 5 years of trajectory. We illustrated multiple non-trivial real patient examples of complex HF trajectories, indicating a promising path for creating highly accurate and scalable longitudinal deep learning models for modeling the chronic disease.</p></div>]]></description>
            <pubDate><![CDATA[2021-01-06T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Prediction of implementing ISO 14031 guidelines using a multilayer perceptron neural network approach]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765766565913-c8f35ba3-4521-40e1-ac53-1d0e6ee10d1c/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244029</link>
            <description><![CDATA[<p class="para" id="N65539">The purpose of this study was to model the link between the implementation of ISO 14031 and ISO 14001. This study connects ISO 14031’s guidelines as independent variables to a dependent variable expressed by the ISO 14001 certification situation of industrial organizations based on the judgments of environmental managers in Saudi Arabia. Applying the quantitative approach using a survey with 596 responses from organizations functioning in 30 economic activities, a multi-layered neural network was trained to examine the relationship between standards and predict whether the organization is ISO 14001 certified in addition to testing the developed network on a group of collected cases. The results demonstrated the ability of the network to classify the organization’s certification status by 94.00% according to the training sample and its ability to predict 91.00% of the test sample, with an overall prediction efficiency of 91.30%. This work provides insights and adds to the environmental performance evaluation literature providing a neural network model based on ISO 14031 guidelines that can be extended to include other international standards such as ISO 9001. This study supports the merging of ISO 14001 with ISO 14031 into a binding standard.</p>]]></description>
            <pubDate><![CDATA[2021-01-06T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Enhancing fine-grained intra-urban dengue forecasting by integrating spatial interactions of human movements between urban regions]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765762220422-ee023446-f6b1-4f59-8a17-0e4247efddf4/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pntd.0008924</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Background</h3><p class="para" id="N65543">As a mosquito-borne infectious disease, dengue fever (DF) has spread through tropical and subtropical regions worldwide in recent decades. Dengue forecasting is essential for enhancing the effectiveness of preventive measures. Current studies have been primarily conducted at national, sub-national, and city levels, while an intra-urban dengue forecasting at a fine spatial resolution still remains a challenging feat. As viruses spread rapidly because of a highly dynamic population flow, integrating spatial interactions of human movements between regions would be potentially beneficial for intra-urban dengue forecasting.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Methodology</h3><p class="para" id="N65549">In this study, a new framework for enhancing intra-urban dengue forecasting was developed by integrating the spatial interactions between urban regions. First, a graph-embedding technique called Node2Vec was employed to learn the embeddings (in the form of an <i>N</i>-dimensional real-valued vector) of the regions from their population flow network. As strongly interacting regions would have more similar embeddings, the embeddings can serve as “interaction features.” Then, the interaction features were combined with those commonly used features (e.g., temperature, rainfall, and population) to enhance the supervised learning–based dengue forecasting models at a fine-grained intra-urban scale.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Results</h3><p class="para" id="N65558">The performance of forecasting models (i.e., SVM, LASSO, and ANN) integrated with and without interaction features was tested and compared on township-level dengue forecasting in Guangzhou, the most threatened sub-tropical city in China. Results showed that models using both common and interaction features can achieve better performance than that using common features alone.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Conclusions</h3><p class="para" id="N65564">The proposed approach for incorporating spatial interactions of human movements using graph-embedding technique is effective, which can help enhance fine-grained intra-urban dengue forecasting.</p></div><p class="para" id="N65542">Dengue fever, a mosquito-borne infectious disease, has become a serious public health problem in many tropical and subtropical regions worldwide, such as Southeast Asian countries and the Guangdong Province in China. In the absence of an effective vaccine at present, disease surveillance and mosquito control remain the primary means of controlling the spread of the disease. At an intra-urban setting, it is important to predict the spatial distribution of future patients, which can help government agencies to establish precise and targeted prevention measures beforehand. Considering the fast virus spread within a city because of a highly dynamic population flow, we proposed a novel approach to enhancing fine-grained intra-urban dengue forecasting by integrating spatial interactions of human movements between urban regions. First, using a graph-embedding model called Node2Vec, the embeddings of the regions were learned from their population interaction network so that strongly interacted regions would have more similar embeddings. Secondly, serving as interaction features, the embeddings were combined with the commonly used features as inputs of the forecasting models. The experimental results indicated that the performance of the models can be improved by incorporating the interaction features, confirming the effectiveness of our proposed strategy in enhancing fine-grained intra-urban dengue forecasting.</p>]]></description>
            <pubDate><![CDATA[2020-12-21T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Rapid detection of fast innovation under the pressure of COVID-19]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765758735830-654c388f-5237-4dcd-a5c8-d8ac3430f173/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244175</link>
            <description><![CDATA[<p class="para" id="N65539">Covid-19 has rapidly redefined the agenda of technological research and development both for academics and practitioners. If the medical scientific publication system has promptly reacted to this new situation, other domains, particularly in new technologies, struggle to map what is happening in their contexts. The pandemic has created the need for a rapid detection of technological convergence phenomena, but at the same time it has made clear that this task is impossible on the basis of traditional patent and publication indicators. This paper presents a novel methodology to perform a rapid detection of the fast technological convergence phenomenon that is occurring under the pressure of the Covid-19 pandemic. The fast detection has been performed thanks to the use of a novel source: the online blogging platform Medium. We demonstrate that the hybrid structure of this social journalism platform allows a rapid detection of innovation phenomena, unlike other traditional sources. The technological convergence phenomenon has been modelled through a network-based approach, analysing the differences of networks computed during two time periods (pre and post COVID-19). The results led us to discuss the repurposing of technologies regarding “Remote Control”, “Remote Working”, “Health” and “Remote Learning”.</p>]]></description>
            <pubDate><![CDATA[2020-12-31T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Application of machine learning in the diagnosis of gastric cancer based on noninvasive characteristics]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765758499345-b686a3c6-4209-4b1e-a5a8-7ff21b449bf5/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244869</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Background</h3><p class="para" id="N65543">The diagnosis of gastric cancer mainly relies on endoscopy, which is invasive and costly. The aim of this study is to develop a predictive model for the diagnosis of gastric cancer based on noninvasive characteristics.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Aims</h3><p class="para" id="N65549">To construct a predictive model for the diagnosis of gastric cancer with high accuracy based on noninvasive characteristics.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Methods</h3><p class="para" id="N65555">A retrospective study of 709 patients at Zhejiang Provincial People's Hospital was conducted. Variables of age, gender, blood cell count, liver function, kidney function, blood lipids, tumor markers and pathological results were analyzed. We used gradient boosting decision tree (GBDT), a type of machine learning method, to construct a predictive model for the diagnosis of gastric cancer and evaluate the accuracy of the model.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Results</h3><p class="para" id="N65561">Of the 709 patients, 398 were diagnosed with gastric cancer; 311 were health people or diagnosed with benign gastric disease. Multivariate analysis showed that gender, age, neutrophil lymphocyte ratio, hemoglobin, albumin, carcinoembryonic antigen (CEA), carbohydrate antigen 125 (CA125) and carbohydrate antigen 199 (CA199) were independent characteristics associated with gastric cancer. We constructed a predictive model using GBDT, and the area under the receiver operating characteristic curve (AUC) of the model was 91%. For the test dataset, sensitivity was 87.0% and specificity 84.1% at the optimal threshold value of 0.56. The overall accuracy was 83.0%. Positive and negative predictive values were 83.0% and 87.8%, respectively.</p></div><div class="section" id="sec005"><h3 class="BHead" id="nov000-5">Conclusion</h3><p class="para" id="N65567">We construct a predictive model to diagnose gastric cancer with high sensitivity and specificity. The model is noninvasive and may reduce the medical cost.</p></div>]]></description>
            <pubDate><![CDATA[2020-12-31T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Predicting self-harm within six months after initial presentation to youth mental health services: A machine learning study]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765758481025-b1fe330d-9e91-4477-91a1-cddf5c6d0179/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0243467</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Background</h3><p class="para" id="N65543">A priority for health services is to reduce self-harm in young people. Predicting self-harm is challenging due to their rarity and complexity, however this does not preclude the utility of prediction models to improve decision-making regarding a service response in terms of more detailed assessments and/or intervention. The aim of this study was to predict self-harm within six-months after initial presentation.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Method</h3><p class="para" id="N65549">The study included 1962 young people (12–30 years) presenting to youth mental health services in Australia. Six machine learning algorithms were trained and tested with ten repeats of ten-fold cross-validation. The net benefit of these models were evaluated using decision curve analysis.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Results</h3><p class="para" id="N65555">Out of 1962 young people, 320 (16%) engaged in self-harm in the six months after first assessment and 1642 (84%) did not. The top 25% of young people as ranked by mean predicted probability accounted for 51.6% - 56.2% of all who engaged in self-harm. By the top 50%, this increased to 82.1%-84.4%. Models demonstrated fair overall prediction (AUROCs; 0.744–0.755) and calibration which indicates that predicted probabilities were close to the true probabilities (brier scores; 0.185–0.196). The net benefit of these models were positive and superior to the ‘treat everyone’ strategy. The strongest predictors were (in ranked order); a history of self-harm, age, social and occupational functioning, sex, bipolar disorder, psychosis-like experiences, treatment with antipsychotics, and a history of suicide ideation.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Conclusion</h3><p class="para" id="N65561">Prediction models for self-harm may have utility to identify a large sub population who would benefit from further assessment and targeted (low intensity) interventions. Such models could enhance health service approaches to identify and reduce self-harm, a considerable source of distress, morbidity, ongoing health care utilisation and mortality.</p></div>]]></description>
            <pubDate><![CDATA[2020-12-31T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Comparison of methods for texture analysis of QUS parametric images in the characterization of breast lesions]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765756790750-e51ab5a9-cbcc-4605-8fa9-86c6de6a38a4/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244965</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Purpose</h3><p class="para" id="N65543">Accurate and timely diagnosis of breast carcinoma is very crucial because of its high incidence and high morbidity. Screening can improve overall prognosis by detecting the disease early. Biopsy remains as the gold standard for pathological confirmation of malignancy and tumour grading. The development of diagnostic imaging techniques as an alternative for the rapid and accurate characterization of breast masses is necessitated. Quantitative ultrasound (QUS) spectroscopy is a modality well suited for this purpose. This study was carried out to evaluate different texture analysis methods applied on QUS spectral parametric images for the characterization of breast lesions.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Methods</h3><p class="para" id="N65549">Parametric images of mid-band-fit (MBF), spectral-slope (SS), spectral-intercept (SI), average scatterer diameter (ASD), and average acoustic concentration (AAC) were determined using QUS spectroscopy from 193 patients with breast lesions. Texture methods were used to quantify heterogeneities of the parametric images. Three statistical-based approaches for texture analysis that include Gray Level Co-occurrence Matrix (GLCM), Gray Level Run-length Matrix (GRLM), and Gray Level Size Zone Matrix (GLSZM) methods were evaluated. QUS and texture-parameters were determined from both tumour core and a 5-mm tumour margin and were used in comparison to histopathological analysis in order to classify breast lesions as either benign or malignant. We developed a diagnostic model using different classification algorithms including linear discriminant analysis (LDA), <i>k</i>-nearest neighbours (KNN), support vector machine with radial basis function kernel (SVM-RBF), and an artificial neural network (ANN). Model performance was evaluated using leave-one-out cross-validation (LOOCV) and hold-out validation.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Results</h3><p class="para" id="N65558">Classifier performances ranged from 73% to 91% in terms of accuracy dependent on tumour margin inclusion and classifier methodology. Utilizing information from tumour core alone, the ANN achieved the best classification performance of 93% sensitivity, 88% specificity, 91% accuracy, 0.95 AUC using QUS parameters and their GLSZM texture features.</p></div><div class="section" id="sec004"><h3 class="BHead" id="nov000-4">Conclusions</h3><p class="para" id="N65564">A QUS-based framework and texture analysis methods enabled classification of breast lesions with &gt;90% accuracy. The results suggest that optimizing method for extracting discriminative textural features from QUS spectral parametric images can improve classification performance. Evaluation of the proposed technique on a larger cohort of patients with proper validation technique demonstrated the robustness and generalization of the approach.</p></div>]]></description>
            <pubDate><![CDATA[2020-12-31T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Imbalanced learning: Improving classification of diabetic neuropathy from magnetic resonance imaging]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765745246497-8bcc11b9-dcee-4ab4-8f5f-5c5ebb721821/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0243907</link>
            <description><![CDATA[<p class="para" id="N65539">One of the fundamental challenges when dealing with medical imaging datasets is class imbalance. Class imbalance happens where an instance in the class of interest is relatively low, when compared to the rest of the data. This study aims to apply oversampling strategies in an attempt to balance the classes and improve classification performance. We evaluated four different classifiers from k-nearest neighbors (k-NN), support vector machine (SVM), multilayer perceptron (MLP) and decision trees (DT) with 73 oversampling strategies. In this work, we used imbalanced learning oversampling techniques to improve classification in datasets that are distinctively sparser and clustered. This work reports the best oversampling and classifier combinations and concludes that the usage of oversampling methods always outperforms no oversampling strategies hence improving the classification results.</p>]]></description>
            <pubDate><![CDATA[2020-12-15T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Optimised genetic algorithm-extreme learning machine approach for automatic COVID-19 detection]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765745092484-947784f4-458d-4604-bbed-47f310120fb2/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0242899</link>
            <description><![CDATA[<p class="para" id="N65539">The coronavirus disease (COVID-19), is an ongoing global pandemic caused by severe acute respiratory syndrome. Chest Computed Tomography (CT) is an effective method for detecting lung illnesses, including COVID-19. However, the CT scan is expensive and time-consuming. Therefore, this work focus on detecting COVID-19 using chest X-ray images because it is widely available, faster, and cheaper than CT scan. Many machine learning approaches such as Deep Learning, Neural Network, and Support Vector Machine; have used X-ray for detecting the COVID-19. Although the performance of those approaches is acceptable in terms of accuracy, however, they require high computational time and more memory space. Therefore, this work employs an Optimised Genetic Algorithm-Extreme Learning Machine (OGA-ELM) with three selection criteria (i.e., random, K-tournament, and roulette wheel) to detect COVID-19 using X-ray images. The most crucial strength factors of the Extreme Learning Machine (ELM) are: (i) high capability of the ELM in avoiding overfitting; (ii) its usability on binary and multi-type classifiers; and (iii) ELM could work as a kernel-based support vector machine with a structure of a neural network. These advantages make the ELM efficient in achieving an excellent learning performance. ELMs have successfully been applied in many domains, including medical domains such as breast cancer detection, pathological brain detection, and ductal carcinoma in situ detection, but not yet tested on detecting COVID-19. Hence, this work aims to identify the effectiveness of employing OGA-ELM in detecting COVID-19 using chest X-ray images. In order to reduce the dimensionality of a histogram oriented gradient features, we use principal component analysis. The performance of OGA-ELM is evaluated on a benchmark dataset containing 188 chest X-ray images with two classes: a healthy and a COVID-19 infected. The experimental result shows that the OGA-ELM achieves 100.00% accuracy with fast computation time. This demonstrates that OGA-ELM is an efficient method for COVID-19 detecting using chest X-ray images.</p>]]></description>
            <pubDate><![CDATA[2020-12-15T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Directions in abusive language training data, a systematic review: Garbage in, garbage out]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765744979881-946c3fe9-08a1-4dc2-9f0f-48619b87347a/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0243300</link>
            <description><![CDATA[<p class="para" id="N65539">Data-driven and machine learning based approaches for detecting, categorising and measuring abusive content such as hate speech and harassment have gained traction due to their scalability, robustness and increasingly high performance. Making effective detection systems for abusive content relies on having the right training datasets, reflecting a widely accepted mantra in computer science: Garbage In, Garbage Out. However, creating training datasets which are large, varied, theoretically-informed and that minimize biases is difficult, laborious and requires deep expertise. This paper systematically reviews 63 publicly available training datasets which have been created to train abusive language classifiers. It also reports on creation of a dedicated website for cataloguing abusive language data hatespeechdata.com. We discuss the challenges and opportunities of open science in this field, and argue that although more dataset sharing would bring many benefits it also poses social and ethical risks which need careful consideration. Finally, we provide evidence-based recommendations for practitioners creating new abusive content training datasets.</p>]]></description>
            <pubDate><![CDATA[2020-12-28T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Using enriched semantic event chains to model human action prediction based on (minimal) spatial information]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765744741118-1f7f1bf5-1770-4a81-90d3-7d5891c8c335/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0243829</link>
            <description><![CDATA[<p class="para" id="N65539">Predicting other people’s upcoming action is key to successful social interactions. Previous studies have started to disentangle the various sources of information that action observers exploit, including objects, movements, contextual cues and features regarding the acting person’s identity. We here focus on the role of static and dynamic inter-object <i>spatial relations</i> that change during an action. We designed a virtual reality setup and tested recognition speed for ten different manipulation actions. Importantly, all objects had been abstracted by emulating them with cubes such that participants could not infer an action using object information. Instead, participants had to rely only on the limited information that comes from the changes in the spatial relations between the cubes. In spite of these constraints, participants were able to predict actions in, on average, less than 64% of the action’s duration. Furthermore, we employed a computational model, the so-called enriched Semantic Event Chain (eSEC), which incorporates the information of different types of spatial relations: (a) objects’ touching/untouching, (b) static spatial relations between objects and (c) dynamic spatial relations between objects during an action. Assuming the eSEC as an underlying model, we show, using information theoretical analysis, that humans mostly rely on a mixed-cue strategy when predicting actions. Machine-based action prediction is able to produce faster decisions based on individual cues. We argue that human strategy, though slower, may be particularly beneficial for prediction of natural and more complex actions with more variable or partial sources of information. Our findings contribute to the understanding of how individuals afford inferring observed actions’ goals even before full goal accomplishment, and may open new avenues for building robots for conflict-free human-robot cooperation.</p>]]></description>
            <pubDate><![CDATA[2020-12-28T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Development of a hybrid model for a partially known intracellular signaling pathway through correction term estimation and neural network modeling]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765744657373-cb3d6aa7-d488-4862-9259-2ddbf464dfc8/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pcbi.1008472</link>
            <description><![CDATA[<p class="para" id="N65539">Developing an accurate first-principle model is an important step in employing systems biology approaches to analyze an intracellular signaling pathway. However, an accurate first-principle model is difficult to be developed since it requires in-depth mechanistic understandings of the signaling pathway. Since underlying mechanisms such as the reaction network structure are not fully understood, significant discrepancy exists between predicted and actual signaling dynamics. Motivated by these considerations, this work proposes a hybrid modeling approach that combines a first-principle model and an artificial neural network (ANN) model so that predictions of the hybrid model surpass those of the original model. First, the proposed approach determines an optimal subset of model states whose dynamics should be corrected by the ANN by examining the correlation between each state and outputs through relative order. Second, an L2-regularized least-squares problem is solved to infer values of the correction terms that are necessary to minimize the discrepancy between the model predictions and available measurements. Third, an ANN is developed to generalize relationships between the values of the correction terms and the system dynamics. Lastly, the original first-principle model is coupled with the developed ANN to finalize the hybrid model development so that the model will possess generalized prediction capabilities while retaining the model interpretability. We have successfully validated the proposed methodology with two case studies, simplified apoptosis and lipopolysaccharide-induced NF<i>κ</i>B signaling pathways, to develop hybrid models with <i>in silico</i> and <i>in vitro</i> measurements, respectively.</p><p class="para" id="N65542">An intracellular signaling pathway is often represented by a set of nonlinear ordinary differential equations, which translate our current knowledge about the signaling pathway into a testable mathematical model. However, predictions from such models are often subject to high uncertainty since many signaling pathways are only partially known beforehand. In this study, we propose a systematic approach to develop a hybrid model to improve model accuracy by combining machine learning and the first-principle modeling. Specifically, model correction terms are learned from discrepancy between model predictions and measurements, and these terms are added to the first-principle model to enhance the prediction accuracy. Once these correction terms are learned from the data, an artificial neural network (ANN) model is developed to find an empirical relation between the model and the correction terms so that the developed ANN can be used to posses improved predictive capabilities even in new operating conditions (i.e., generalizability). The final hybrid model is then constructed by coupling the first-principle model with the developed ANN.</p>]]></description>
            <pubDate><![CDATA[2020-12-14T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Evidence of vascular endothelial dysfunction in Wooden Breast disorder in chickens: Insights through gene expression analysis, ultra-structural evaluation and supervised machine learning methods]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740858598-835b1cfe-b38c-42e5-988f-c14881a47a9f/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0243983</link>
            <description><![CDATA[<p class="para" id="N65539">Several gene expression studies have been previously conducted to characterize molecular basis of Wooden Breast myopathy in commercial broiler chickens. These studies have generally used a limited sample size and relied on a binary disease outcome (unaffected or affected by Wooden Breast), which are appropriate for an initial investigation. However, to identify biomarkers of disease severity and development, it is necessary to use a large number of samples with a varying degree of disease severity. Therefore, in this study, we assayed a relatively large number of samples (n = 96) harvested from the <i>pectoralis major</i> muscle of unaffected (U), partially affected (P) and markedly affected (A) chickens. Gene expression analysis was conducted using the nCounter MAX Analysis System and data were analyzed using four different supervised machine-learning methods, including support vector machines (SVM), random forests (RF), elastic net logistic regression (ENET) and Lasso logistic regression (LASSO). The SVM method achieved the highest prediction accuracy for both three-class (U, P and A) and two-class (U and P+A) classifications with 94% prediction accuracy for two-class classification and 85% for three-class classification. The results also identified biomarkers of Wooden Breast severity and development. Additionally, gene expression analysis and ultrastructural evaluations provided evidence of vascular endothelial cell dysfunction in the early pathogenesis of Wooden Breast.</p>]]></description>
            <pubDate><![CDATA[2021-01-04T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[EXSEQREG: Explaining sequence-based NLP tasks with regions with a case study using morphological features for named entity recognition]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740573879-f550ec87-cb6d-4d8a-b227-a8eed1280acd/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244179</link>
            <description><![CDATA[<p class="para" id="N65539">The state-of-the-art systems for most natural language engineering tasks employ machine learning methods. Despite the improved performances of these systems, there is a lack of established methods for assessing the quality of their predictions. This work introduces a method for explaining the predictions of any sequence-based natural language processing (NLP) task implemented with any model, neural or non-neural. Our method named EXSEQREG introduces the concept of region that links the prediction and features that are potentially important for the model. A region is a list of positions in the input sentence associated with a single prediction. Many NLP tasks are compatible with the proposed explanation method as regions can be formed according to the nature of the task. The method models the prediction probability differences that are induced by careful removal of features used by the model. The output of the method is a list of importance values. Each value signifies the impact of the corresponding feature on the prediction. The proposed method is demonstrated with a neural network based named entity recognition (NER) tagger using Turkish and Finnish datasets. A qualitative analysis of the explanations is presented. The results are validated with a procedure based on the mutual information score of each feature. We show that this method produces reasonable explanations and may be used for i) assessing the degree of the contribution of features regarding a specific prediction of the model, ii) exploring the features that played a significant role for a trained model when analyzed across the corpus.</p>]]></description>
            <pubDate><![CDATA[2020-12-30T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[High throughput mathematical modeling and multi-objective evolutionary algorithms for plant tissue culture media formulation: Case study of pear rootstocks]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765740404798-cc5f1827-a384-4d31-bf4f-819b554a1794/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0243940</link>
            <description><![CDATA[<p class="para" id="N65539">Simplified prediction of the interactions of plant tissue culture media components is of critical importance to efficient development and optimization of new media. We applied two algorithms, gene expression programming (GEP) and M5’ model tree, to predict the effects of media components on in vitro proliferation rate (PR), shoot length (SL), shoot tip necrosis (STN), vitrification (Vitri) and quality index (QI) in pear rootstocks (Pyrodwarf and OHF 69). In order to optimize the selected prediction models, as well as achieving a precise multi-optimization method, multi-objective evolutionary optimization algorithms using genetic algorithm (GA) and particle swarm optimization (PSO) techniques were compared to the mono-objective GA optimization technique. A Gamma test (GT) was used to find the most important determinant input for optimizing each output factor. GEP had a higher prediction accuracy than M5’ model tree. GT results showed that BA (Γ = 4.0178), Mesos (Γ = 0.5482), Mesos (Γ = 184.0100), Micros (Γ = 136.6100) and Mesos (Γ = 1.1146), for PR, SL, STN, Vitri and QI respectively, were the most important factors in culturing OHF 69, while for Pyrodwarf culture, BA (Γ = 10.2920), Micros (Γ = 0.7874), NH<sub>4</sub>NO<sub>3</sub> (Γ = 166.410), KNO<sub>3</sub> (Γ = 168.4400), and Mesos (Γ = 1.4860) were the most important influences on PR, SL, STN, Vitri and QI respectively. The PSO optimized GEP models produced the best outputs for both rootstocks.</p>]]></description>
            <pubDate><![CDATA[2020-12-18T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Evolutionary model discovery of causal factors behind the socio-agricultural behavior of the Ancestral Pueblo]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765739857699-2a3d190e-2e49-412c-894a-d86e939cab3f/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0239922</link>
            <description><![CDATA[<p class="para" id="N65539">Agent-based modeling of artificial societies allows for the validation and analysis of human-interpretable, causal explanations of human behavior that generate society-scale phenomena. However, parameter calibration is insufficient to conduct data-driven explorations that are adequate in evaluating the importance of causal factors that constitute agent rules that match real-world individual-scale generative behaviors. We introduce evolutionary model discovery, a framework that combines genetic programming and random forest regression to evaluate the importance of a set of causal factors hypothesized to affect the individual’s decision-making process. With evolutionary model discovery, we investigated the farm plot seeking behavior of the Ancestral Pueblo of the Long House Valley simulated in the Artificial Anasazi model. We evaluated the importance of causal factors unconsidered in the original model, which we hypothesized to have affected the decision-making process. Our findings, concur with other archaeological studies on the Ancestral Pueblo communities during the Pueblo II period, which indicate the existence of cross-village polities, hierarchical organization, and dependence on the viability of the agricultural niche. Contrary to the original Artificial Anasazi model, where closeness was the sole factor driving farm plot selection, selection of higher quality land, distancing from failed farm plots, and desire for social presence are found to be more important. Finally, models updated with farm selection strategies designed by incorporating these insights showed significant improvements in accuracy and robustness over the original Artificial Anasazi model.</p>]]></description>
            <pubDate><![CDATA[2020-12-18T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Validation of motion tracking as tool for observational toothbrushing studies]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765738961948-d1e35c08-9d86-4bce-be58-e79ec5129991/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244678</link>
            <description><![CDATA[<p class="para" id="N65539">Video observation (VO) is an established tool for observing toothbrushing behaviour, however, it is a subjective method requiring thorough calibration and training, and the toothbrush position is not always clearly visible. As automated tracking of motions may overcome these disadvantages, the study aimed to compare observational data of habitual toothbrushing as well as of post-instruction toothbrushing obtained from motion tracking (MT) to observational data obtained from VO. One-hundred-three subjects (37.4±14.7 years) were included and brushed their teeth with a manual (MB; n = 51) or a powered toothbrush (PB; n = 52) while being simultaneously video-filmed and tracked. Forty-six subjects were then instructed how to brush their teeth systematically and were filmed/tracked for a second time. Videos were analysed with INTERACT (Mangold, Germany); parameters of interest were toothbrush position, brushing time, changes between areas (events) and the Toothbrushing Systematic Index (TSI). Overall, the median proportion (min; max) of identically classified toothbrush positions (both sextant/surface correct) in a brushing session was 87.8% (50.0; 96.9), which was slightly higher for MB compared to PB (90.3 (50.0; 96.9) vs 86.5 (63.7; 96.5) resp.; p = 0.005). The number of events obtained from MT was higher than from VO (p &lt; 0.001) with a moderate to high correlation between them (MB: <i>ρ</i> = 0.52, p &lt; 0.001; PB: <i>ρ</i> = 0.87; p &lt; 0.001). After instruction, both methods revealed a significant increase of the TSI regardless of the toothbrush type (p &lt; 0.001 each). Motion tracking is a suitable tool for observing toothbrushing behaviour, is able to measure improvements after instruction, and can be used with both manual and powered toothbrushes.</p>]]></description>
            <pubDate><![CDATA[2020-12-30T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Distinguishing Discoid and Centripetal Levallois methods through machine learning]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765610217090-ce5618d0-234d-45cf-84cf-a7ec318a25b6/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0244288</link>
            <description><![CDATA[<p class="para" id="N65539">In this paper, we apply Machine Learning (ML) algorithms to study the differences between Discoid and Centripetal Levallois methods. For this purpose, we have used experimentally knapped flint flakes, measuring several parameters that have been analyzed by seven ML algorithms. From these analyses, it has been possible to demonstrate the existence of statistically significant differences between Discoid products and Centripetal Levallois products, thus contributing with new data and a new method to this traditional debate. The new approach enabled differentiating the blanks created by both knapping methods with an accuracy &gt;80% using only ten typometric variables. The most relevant variables were maximum length, width to the 25%, 50% and 75% of the flake length, external and internal platform angles, maximum width and number of dorsal scars. This study also demonstrates the advantages of the application of multivariate ML methods to lithic studies.</p>]]></description>
            <pubDate><![CDATA[2020-12-23T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Causal relations of health indices inferred statistically using the DirectLiNGAM algorithm from big data of Osaka prefecture health checkups]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608445246-164ff459-239b-476e-920d-9dc58695a511/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0243229</link>
            <description><![CDATA[<p class="para" id="N65539">Causal relations among many statistical variables have been assessed using a Linear non-Gaussian Acyclic Model (LiNGAM). Using access to large amounts of health checkup data from Osaka prefecture obtained during the six fiscal years of years 2012–2017, we applied the DirectLiNGAM algorithm as a trial to extract causal relations among health indices for age groups and genders. Results show that LiNGAM yields interesting and reasonable results, suggesting causal relations and correlation among the statistical indices used for these analyses.</p>]]></description>
            <pubDate><![CDATA[2020-12-23T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[MeshCut data augmentation for deep learning in computer vision]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765608417501-2c8cbccd-7b3b-4be4-be90-40007adab643/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pone.0243613</link>
            <description><![CDATA[<p class="para" id="N65539">To solve overfitting in machine learning, we propose a novel data augmentation method called MeshCut, which uses a mesh-like mask to segment the whole image to achieve more partial diversified information. In our experiments, this strategy outperformed the existing augmentation strategies and achieved state-of-the-art results in a variety of computer vision tasks. MeshCut is also an easy-to-implement strategy that can efficiently improve the performance of the existing convolutional neural network models by a good margin without careful hand-tuning. The performance of such a strategy can be further improved by incorporating it into other augmentation strategies, which can make MeshCut a promising baseline strategy for future data augmentation algorithms.</p>]]></description>
            <pubDate><![CDATA[2020-12-23T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Assessing the risk of dengue severity using demographic information and laboratory test results with machine learning]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1765603913819-574d7888-68a4-4c27-8942-4a7e8443a4f2/cover.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/10.1371/journal.pntd.0008960</link>
            <description><![CDATA[<div class="section" id="sec001"><h3 class="BHead" id="nov000-1">Background</h3><p class="para" id="N65543">Dengue virus causes a wide spectrum of disease, which ranges from subclinical disease to severe dengue shock syndrome. However, estimating the risk of severe outcomes using clinical presentation or laboratory test results for rapid patient triage remains a challenge. Here, we aimed to develop prognostic models for severe dengue using machine learning, according to demographic information and clinical laboratory data of patients with dengue.</p></div><div class="section" id="sec002"><h3 class="BHead" id="nov000-2">Methodology/Principal findings</h3><p class="para" id="N65549">Out of 1,581 patients in the National Cheng Kung University Hospital with suspected dengue infections and subjected to NS1 antigen, IgM and IgG, and qRT-PCR tests, 798 patients including 138 severe cases were enrolled in the study. The primary target outcome was severe dengue. Machine learning models were trained and tested using the patient dataset that included demographic information and qualitative laboratory test results collected on day 1 when they sought medical advice. To develop prognostic models, we applied various machine learning methods, including logistic regression, random forest, gradient boosting machine, support vector classifier, and artificial neural network, and compared the performance of the methods. The artificial neural network showed the highest average discrimination area under the receiver operating characteristic curve (0.8324 ± 0.0268) and balance accuracy (0.7523 ± 0.0273). According to the model explainer that analyzed the contributions/co-contributions of the different factors, patient age and dengue NS1 antigenemia were the two most important risk factors associated with severe dengue. Additionally, co-existence of anti-dengue IgM and IgG in patients with dengue increased the probability of severe dengue.</p></div><div class="section" id="sec003"><h3 class="BHead" id="nov000-3">Conclusions/Significance</h3><p class="para" id="N65555">We developed prognostic models for the prediction of dengue severity in patients, using machine learning. The discriminative ability of the artificial neural network exhibited good performance for severe dengue prognosis. This model could help clinicians obtain a rapid prognosis during dengue outbreaks. However, the model requires further validation using external cohorts in future studies.</p></div><p class="para" id="N65542">Dengue virus infects millions of people annually and is associated with a high mortality rate. When outbreaks occur, hospitals are often overcrowded with patients. Thus, novel approaches are required to accelerate patient triage for hospitalization, or further intensive care. Machine learning is being widely applied for resolving various problems, including medical diagnosis and outcome prediction. Here, we combined information from patients, including age, sex, and rapid virus test results, to develop a machine learning model for severe outcome prediction. The developed machine learning model displayed good performance for severe dengue disease prediction, and all information required for the model could be easily obtained. We also found that patients who were over 60 years old, who had detectable nonstructural protein-1 from dengue virus, or who had both detectable anti-dengue IgM and IgG antibodies in their sera, had a greater risk of progression to severe dengue. This study established a new approach to predict dengue disease outcomes by applying machine learning and defined the risk factors for severity prediction.</p>]]></description>
            <pubDate><![CDATA[2020-12-23T00:00]]></pubDate>
        </item><item>
            <title><![CDATA[Let’s Talk AI]]></title>
            <media:thumbnail url="https://storage.googleapis.com/nova-demo-unsecured-files/unsecured/content-1764957438770-fef54231-6b7a-4a00-9a28-9627a6bb7dc0/9783032090089.png"></media:thumbnail>
            <link>https://www.novareader.co/book/isbn/9783032090089</link>
            <description><![CDATA[This unique open access volume represents an interdisciplinary dialog with top-class researchers and practitioners on the Artificial Intelligence revolution. The contributions derive from structured interviews with community leaders who examine recent developments in AI and offer predictions about its impact on science, technology, business, and society. The interviewees, leading thinkers and practitioners from fields such as Computer Science, Software Engineering, Philosophy, Psychology, and Law, gathered at the multidisciplinary event AISoLA, where they discussed developments in AI technologies, tools, and research, and these exchanges were shaped into this coherent survey.

In the debates surrounding the rapid advancement of AI it’s clear that while there is much sharing of ideas, in fact true cooperation and alignment is rare, most fields still operate in silos. Given the pervasive nature of AI, it’s critical that we move forward carefully to ensure that our decisions and progress are well-considered. This book serves as a call to action, reminding us of our collective responsibility to shape the future. The authors offer points of agreement and contrast, the reader will better understand how we may still maintain control over technological progress in AI and its societal impact.]]></description>
            <pubDate><![CDATA[2025-10-26T18:30]]></pubDate>
        </item>
    </channel>
</rss>