Papers.
Research connected to its authors, projects, companies, talks, events, and the rest of the graph.
Add a paper ↗Lexical and acoustic speech features relating to Alzheimer’s disease pathology
medRxiv · DOI 10.1101/2021.09.27.21264148 · 2 citations · Source: openalex+first-party-career-authorityMark Liberman, Sunghye Cho, Katheryn A Q Cousins, Sanjana Shellikeri, Sharon Ash, David J. Irwin, Murray Grossman, Naomi Nevler · 8 authors totalLaunching a saliva-based SARS-CoV-2 surveillance testing program on a university campus
medRxiv · DOI 10.1101/2021.01.24.21250385 · 6 citations · Source: openalex+authoritative-profilePetros Giannikopoulos, Alexander J. Ehrenberg, Erica A. Moehle, Cara E. Brook, Andrew H. Doudna Cate, Lea B. Witkowsky, Rohan Sachdeva, Ariana Hirsh · 45 authors totalRobotic RNA extraction for SARS-CoV-2 surveillance using saliva samples
medRxiv · DOI 10.1101/2021.01.10.21249151 · 9 citations · Source: openalex+authoritative-profilePetros Giannikopoulos, Jennifer Hamilton, Elizabeth C. Stahl, Connor A. Tsuchida, Enrique Lin-Shiao, C. Kimberly Tsui, Kathleen Pestal, Holly K. Gildea · 21 authors totalEpitome: predicting epigenetic events in novel cell types with multi-cell deep ensemble learning
Nucleic Acids Research · DOI 10.1093/nar/gkab676 · 4 citations · Source: semantic-scholarThe accumulation of large epigenomics data consortiums provides us with the opportunity to extrapolate existing knowledge to new cell types and conditions. We propose Epitome, a deep neural network that learns similarities of chromatin accessibility between well characterized reference cell types and a query cellular context, and copies over signal of transcription factor binding and modification of histones from reference cell types when chromatin profiles are similar to the query. Epitome achieves state-of-the-art accuracy when predicting transcription factor binding sites on novel cellular contexts, and can further improve predictions as more epigenetic signals are collected from both reference cell types and the query cellular context of interest.
Alyssa Morrow, J. Weston Hughes, Jahnavi Singh, Anthony D. Joseph, Nir Yosef · 5 authors totalThe emotional and mental health impact of the murder of George Floyd on the US population
Proceedings of the National Academy of Sciences · DOI 10.1073/pnas.2109139118 · 225 citations · Source: openalex= 319,471). According to the Gallup data, in the week following Floyd's death, anger and sadness increased to unprecedented levels in the US population. During this period, more than a third of the US population reported these emotions. These increases were more pronounced for Black Americans, nearly half of whom reported these emotions. According to the US Census Household Pulse data, in the week following Floyd's death, depression and anxiety severity increased among Black Americans at significantly higher rates than that of White Americans. Our estimates suggest that this increase corresponds to an additional 900,000 Black Americans who would have screened positive for depression, associated with a burden of roughly 2.7 million to 6.3 million mentally unhealthy days.
Lyle Ungar, Johannes C. Eichstaedt, Garrick Sherman, Salvatore Giorgi, Steven O. Roberts, Megan E. Reynolds, Sharath Chandra Guntuku · 7 authors totalAn evidence review of face masks against COVID-19
PNAS · DOI 10.1073/pnas.2014564118 · 1,015 citations · Source: semantic-scholarThe science around the use of masks by the public to impede COVID-19 transmission is advancing rapidly. In this narrative review, we develop an analytical framework to examine mask usage, synthesizing the relevant literature to inform multiple areas: population impact, transmission characteristics, source control, wearer protection, sociological considerations, and implementation considerations. A primary route of transmission of COVID-19 is via respiratory particles, and it is known to be transmissible from presymptomatic, paucisymptomatic, and asymptomatic individuals. Reducing disease spread requires two things: limiting contacts of infected individuals via physical distancing and other measures and reducing the transmission probability per contact. The preponderance of evidence indicates that mask wearing reduces transmissibility per contact by reducing transmission of infected respiratory particles in both laboratory and clinical contexts. Public mask wearing is most effective at reducing spread of the virus when compliance is high. Given the current shortages of medical masks, we recommend the adoption of public cloth mask wearing, as an effective form of source control, in conjunction with existing hygiene, distancing, and contact tracing strategies. Because many respiratory particles become smaller due to evaporation, we recommend increasing focus on a previously overlooked aspect of mask usage: mask wearing by infectious people (“source control”) with benefits at the population level, rather than only mask wearing by susceptible people, such as health care workers, with focus on individual outcomes. We recommend that public officials and governments strongly encourage the use of widespread face masks in public, including the use of appropriate regulation.
Jeremy Howard, J. Howard, Austin Huang, Zhiyuan Li, Zeynep Tufekci, V. Ždímal, H. van der Westhuizen, A. von Delft · 19 authors totalLexical and Acoustic Characteristics of Young and Older Healthy Adults
Journal of Speech Language and Hearing Research · DOI 10.1044/2020_jslhr-19-00384 · 26 citations · Source: openalex+first-party-career-authorityMark Liberman, Sunghye Cho, Naomi Nevler, Sanjana Shellikeri, Natalia Parjane, David J. Irwin, Neville Ryant, Sharon Ash · 10 authors totalEstimation of continuous valence and arousal levels from faces in naturalistic conditions
Nature Machine Intelligence · DOI 10.1038/s42256-020-00280-0 · 260 citations · Source: semantic-scholarFacial affect analysis aims to create new types of human–computer interactions by enabling computers to better understand a person’s emotional state in order to provide ad hoc help and interactions. Since discrete emotional classes (such as anger, happiness, sadness and so on) are not representative of the full spectrum of emotions displayed by humans on a daily basis, psychologists typically rely on dimensional measures, namely valence (how positive the emotional display is) and arousal (how calming or exciting the emotional display looks like). However, while estimating these values from a face is natural for humans, it is extremely difficult for computer-based systems and automatic estimation of valence and arousal in naturalistic conditions is an open problem. Additionally, the subjectivity of these measures makes it hard to obtain good quality data. Here we introduce a novel deep neural network architecture to analyse facial affect in naturalistic conditions with a high level of accuracy. The proposed network integrates face alignment and jointly estimates both categorical and continuous emotions in a single pass, making it suitable for real-time applications. We test our method on three challenging datasets collected in naturalistic conditions and show that our approach outperforms all previous methods. We also discuss caveats regarding the use of this tool, and ethical aspects that must be considered in its application. The annotation of the visual signs of emotions ca
Jean Kossaifi, Antoine Toisoul, Adrian Bulat, Georgios Tzimiropoulos, M. Pantic · 5 authors totalDeep learning-enabled medical computer vision
npj Digital Medicine · DOI 10.1038/s41746-020-00376-2 · 1,224 citations · Source: semantic-scholarA decade of unprecedented progress in artificial intelligence (AI) has demonstrated the potential for many fields—including medicine—to benefit from the insights that AI techniques can extract from data. Here we survey recent progress in the development of modern computer vision techniques—powered by deep learning—for medical applications, focusing on medical imaging, medical video, and clinical deployment. We start by briefly summarizing a decade of progress in convolutional neural networks, including the vision tasks they enable, in the context of healthcare. Next, we discuss several example medical imaging applications that stand to benefit—including cardiology, pathology, dermatology, ophthalmology–and propose new avenues for continued work. We then expand into general medical video, highlighting ways in which clinical workflows can integrate computer vision to enhance care. Finally, we discuss the challenges and hurdles required for real-world clinical deployment of these technologies.
Richard Socher, A. Esteva, Katherine Chou, Serena Yeung, N. Naik, Ali Madani, A. Mottaghi, Yun Liu · 10 authors totalDeep Learning-Based Point-Scanning Super-Resolution Imaging
Nature Methods · DOI 10.1038/s41592-021-01080-z · 183 citations · Source: semantic-scholarPoint-scanning imaging systems are among the most widely used tools for high-resolution cellular and tissue imaging, benefiting from arbitrarily defined pixel sizes. The resolution, speed, sample preservation and signal-to-noise ratio (SNR) of point-scanning systems are difficult to optimize simultaneously. We show these limitations can be mitigated via the use of deep learning-based supersampling of undersampled images acquired on a point-scanning system, which we term point-scanning super-resolution (PSSR) imaging. We designed a ‘crappifier’ that computationally degrades high SNR, high-pixel resolution ground truth images to simulate low SNR, low-resolution counterparts for training PSSR models that can restore real-world undersampled images. For high spatiotemporal resolution fluorescence time-lapse data, we developed a ‘multi-frame’ PSSR approach that uses information in adjacent frames to improve model predictions. PSSR facilitates point-scanning image acquisition with otherwise unattainable resolution, speed and sensitivity. All the training data, models and code for PSSR are publicly available at 3DEM.org. Point-scanning super-resolution imaging uses deep learning to supersample undersampled images and enable time-lapse imaging of subcellular events. An accompanying ‘crappifier’ rapidly generates quality training data for robust performance.
Jeremy Howard, Linjing Fang, Fred Monroe, S. Novak, Lyndsey Kirk, Cara R. Schiavon, S. B. Yu, Tong Zhang · 19 authors totalPublisher Correction: Accelerated RNA detection using tandem CRISPR nucleases
Nature Chemical Biology · DOI 10.1038/s41589-021-00882-8 · 8 citations · Source: openalex+authoritative-profilePetros Giannikopoulos, Tina Y. Liu, Gavin J. Knott, Dylan C. J. Smock, John J. Desmarais, Sungmin Son, Abdul Bhuiya, Shrutee Jakhanwal · 48 authors totalAccelerated RNA detection using tandem CRISPR nucleases
Nature Chemical Biology · DOI 10.1038/s41589-021-00842-2 · 272 citations · Source: openalex+authoritative-profilePetros Giannikopoulos, Tina Y. Liu, Gavin J. Knott, Dylan C. J. Smock, John J. Desmarais, Sungmin Son, Abdul Bhuiya, Shrutee Jakhanwal · 48 authors totalMegastudies improve the impact of applied behavioural science
Nature · DOI 10.1038/s41586-021-04128-4 · 266 citations · Source: openalexLyle Ungar, Katherine L. Milkman, Dena M. Gromet, Hung S. Ho, Joseph Kay, Timothy W. Lee, Pepi Pandiloski, Yeji Park · 43 authors totalSelf-guarding of MORC3 enables virulence factor-triggered immunity
Nature · DOI 10.1038/s41586-021-04054-5 · 73 citations · Source: semantic-scholarPathogens use virulence factors to inhibit the immune system1. The guard hypothesis2,3 postulates that hosts monitor (or ‘guard’) critical innate immune pathways such that their disruption by virulence factors provokes a secondary immune response1. Here we describe a ‘self-guarded’ immune pathway in human monocytes, in which guarding and guarded functions are combined in one protein. We find that this pathway is triggered by ICP0, a key virulence factor of herpes simplex virus type 1, resulting in robust induction of anti-viral type I interferon (IFN). Notably, induction of IFN by ICP0 is independent of canonical immune pathways and the IRF3 and IRF7 transcription factors. A CRISPR screen identified the ICP0 target MORC34 as an essential negative regulator of IFN. Loss of MORC3 recapitulates the IRF3- and IRF7-independent IFN response induced by ICP0. Mechanistically, ICP0 degrades MORC3, which leads to de-repression of a MORC3-regulated DNA element (MRE) adjacent to the IFNB1 locus. The MRE is required in cis for IFNB1 induction by the MORC3 pathway, but is not required for canonical IFN-inducing pathways. As well as repressing the MRE to regulate IFNB1, MORC3 is also a direct restriction factor of HSV-15. Our results thus suggest a model in which the primary anti-viral function of MORC3 is self-guarded by its secondary IFN-repressing function—thus, a virus that degrades MORC3 to avoid its primary anti-viral function will unleash the secondary anti-viral IFN response. MORC3 is revealed as an essential negative regulator of the anti-viral interferon response that functions in an innate immune pathway that detects viral virulence factors.
Alyssa Morrow, Moritz M. Gaidt, Marian R. Fairgrieve, Jonathan P. Karr, Nir Yosef, Russell E. Vance · 6 authors totalNatural language processing methods are sensitive to sub-clinical linguistic differences in schizophrenia spectrum disorders
Schizophrenia · DOI 10.1038/s41537-021-00154-3 · 134 citations · Source: openalex+first-party-career-authorityMark Liberman, Sunny X. Tang, Reno Kriz, Sunghye Cho, Suh Jung Park, Jenna Harowitz, Raquel E. Gur, Mahendra T. Bhati · 10 authors totalClosed- and open-vocabulary approaches to text analysis: A review, quantitative comparison, and recommendations.
Psychological Methods · DOI 10.1037/met0000349 · 173 citations · Source: openalexTechnology now makes it possible to understand efficiently and at large scale how people use language to reveal their everyday thoughts, behaviors, and emotions. Written text has been analyzed through both theory-based, closed-vocabulary methods from the social sciences as well as data-driven, open-vocabulary methods from computer science, but these approaches have not been comprehensively compared. To provide guidance on best practices for automatically analyzing written text, this narrative review and quantitative synthesis compares five predominant closed- and open-vocabulary methods: Linguistic Inquiry and Word Count (LIWC), the General Inquirer, DICTION, Latent Dirichlet Allocation, and Differential Language Analysis. We compare the linguistic features associated with gender, age, and personality across the five methods using an existing dataset of Facebook status updates and self-reported survey data from 65,896 users. Results are fairly consistent across methods. The closed-vocabulary approaches efficiently summarize concepts and are helpful for understanding how people think, with LIWC2015 yielding the strongest, most parsimonious results. Open-vocabulary approaches reveal more specific and concrete patterns across a broad range of content domains, better address ambiguous word senses, and are less prone to misinterpretation, suggesting that they are well-suited for capturing the nuances of everyday psychological processes. We detail several errors that can occur in closed-vocabulary analyses, the impact of sample size, number of words per user and number of topics included in open-vocabulary analyses, and implications of different analytical decisions. We conclude with recommendations for researchers, advocating for a complementary approach that combines...
Lyle Ungar, Johannes C. Eichstaedt, Margaret L. Kern, David B. Yaden, H. Andrew Schwartz, Salvatore Giorgi, Gregory Park, Courtney A. Hagan · 13 authors totalThe #ddj Hashtag on Twitter
DOI 10.1017/9789048542079.037 · 0 citations · Source: openalex+first-party-career-authorityMarc Smith, Eunice Au, Marc A. Smith · 3 authors totalSafe functional systems through integrity types and verified assembly
Theoretical Computer Science · DOI 10.1016/j.tcs.2020.09.039 · 1 citations · Source: semantic-scholar+dblpAbstract Building a trustworthy life-critical embedded system requires deep reasoning about the potential effects that sequences of machine instructions can have on full system operation. Rather than trying to analyze complete binaries and the countless ways their instructions can interact with one another — memory, side effects, control registers, implicit state, etc. — we explore a new approach. We propose an architecture controlled by a thin computational layer designed to tightly correspond with the lambda calculus, drawing on principles of functional programming to bring the assembly much closer to myriad reasoning frameworks, such as the Coq proof assistant. This approach allows assembly-level verified versions of critical code to operate safely in tandem with arbitrary code, including imperative and unverified system components, without the need for large supporting trusted computing bases. We demonstrate that this computational layer can be built in such a way as to simultaneously provide full programmability and compact, precise, and complete semantics, while still using hardware resources comparable to normal embedded systems. To demonstrate the practicality of this approach, our FPGA-implemented prototype runs an embedded medical application which monitors and treats life-threatening arrhythmias. Though the system integrates untrusted and imperative components, our architecture allows for the formal verification of multiple properties of the end-to-end system. We present a proof of correctness of the assembly-level implementation of the core algorithm in Coq, the integrity of trusted data via a non-interference proof, and a guarantee that our prototype meets critical timing requirements.
Jared Roesch, Michael Christensen, Joseph McMahan, L. Nichols, T. Sherwood, Ben Hardekopf · 6 authors totalSpark NLP: Natural Language Understanding at Scale
Software Impacts · DOI 10.1016/j.simpa.2021.100058 · 6 citations · Source: openalexSpark NLP is a Natural Language Processing (NLP) library built on top of Apache Spark ML. It provides simple, performant & accurate NLP annotations for machine learning pipelines that can scale easily in a distributed environment. Spark NLP comes with 1100+ pretrained pipelines and models in more than 192+ languages. It supports nearly all the NLP tasks and modules that can be used seamlessly in a cluster. Downloaded more than 2.7 million times and experiencing 9x growth since January 2020, Spark NLP is used by 54% of healthcare organizations as the world's most widely used NLP library in the enterprise.
David Talby, Veysel Kocaman · 2 authors totalSame data, different conclusions: Radical dispersion in empirical results when independent analysts operationalize and test the same hypothesis
Organizational Behavior and Human Decision Processes · DOI 10.1016/j.obhdp.2021.02.003 · 115 citations · Source: semantic-scholarIn this crowdsourced initiative, independent analysts used the same dataset to test two hypotheses regarding the effects of scientists’ gender and professional status on verbosity during group meetings. Not only the analytic approach but also the operationalizations of key variables were left unconstrained and up to individual analysts. For instance, analysts could choose to operationalize status as job title, institutional ranking, citation counts, or some combination. To maximize transparency regarding the process by which analytic choices are made, the analysts used a platform we developed called DataExplained to justify both preferred and rejected analytic paths in real time. Analyses lacking sufficient detail, reproducible code, or with statistical errors were excluded, resulting in 29 analyses in the final sample. Researchers reported radically different analyses and dispersed empirical outcomes, in a number of cases obtaining significant effects in opposite directions for the same research question. A Boba multiverse analysis demonstrates that decisions about how to operationalize variables explain variability in outcomes above and beyond statistical choices (e.g., covariates). Subjective researcher decisions play a critical role in driving the reported empirical results, underscoring the need for open data, systematic robustness checks, and transparency regarding both analytic paths taken and not taken. Implications for organizations and leaders, whose decision making relies in part on scientific findings, consulting reports, and internal analyses by data scientists, are discussed.
Eduardo Ariño de la Rubia, Martin Schweinsberg, Michael Feldman, Nicola Staub, O. V. D. Akker, R. V. Aert, M. V. Assen, Yang Liu · 179 authors totalFair Top-k Ranking with multiple protected groups
Information Processing & Management · DOI 10.1016/j.ipm.2021.102707 · 57 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Meike Zehlike, Tom Sühr, Ricardo Baeza‐Yates, Francesco Bonchi, Carlos Castillo, Sara Hajian · 7 authors totalB-PO01-081 MACHINE LEARNING OF THE ELECTROCARDIOGRAM IDENTIFIES CARDIAC WALL MOTION ABNORMALITIES BEYOND THE Q WAVE
Heart Rhythm · DOI 10.1016/j.hrthm.2021.06.226 · 0 citations · Source: openalex+authoritative-profilePeter Bailis, Albert J. Rogers, Neal K. Bhatia, James Tooley, Vyom Thakkar, Jessica Torres, Justin Xu, Jagteshwar Tung · 17 authors totalAutomated analysis of lexical features in frontotemporal degeneration
Cortex · DOI 10.1016/j.cortex.2021.01.012 · 43 citations · Source: openalex+first-party-career-authorityMark Liberman, Sunghye Cho, Naomi Nevler, Sharon Ash, Sanjana Shellikeri, David J. Irwin, Lauren Massimo, Katya Rascovsky · 10 authors totalConTrib: Maintaining fairness in decentralized big tech alternatives by accounting work
Computer Networks · DOI 10.1016/j.comnet.2021.108081 · 5 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Martijn de Vos · 2 authors totalUsing Computerized Human Language Technology to Generate Biomarkers for Disturbances in Language
Biological Psychiatry · DOI 10.1016/j.biopsych.2021.02.932 · 0 citations · Source: openalex+first-party-career-authorityMark Liberman, Rony Krell, Wenqing Tang, Katrin Hänsel, Michael Sobolev, Sunghye Cho, Sarah Berretta, Aarush Mehta · 9 authors totalCorrection to: Towards intellectual freedom in an AI Ethics Global Community
AI and Ethics · DOI 10.1007/s43681-021-00059-y · 1 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Christoph Ebell, Ricardo Baeza‐Yates, Richard Benjamins, Hengjin Cai, Mark Coeckelbergh, Tania Duarte, Merve Hickok · 17 authors totalTowards intellectual freedom in an AI Ethics Global Community
AI and Ethics · DOI 10.1007/s43681-021-00052-5 · 37 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Christoph Ebell, Ricardo Baeza‐Yates, Richard Benjamins, Hengjin Cai, Mark Coeckelbergh, Tania Duarte, Merve Hickok · 17 authors totalXChange: A Universal Mechanism for Asset Exchange between Permissioned Blockchains
World Wide Web · DOI 10.1007/s11280-021-00870-x · 11 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Martijn de Vos, Can Umut Ileri · 3 authors totalOpen Vocabulary Object Detection with Pseudo Bounding-Box Labels
European Conference on Computer Vision · DOI 10.1007/978-3-031-20080-9_16 · arXiv 2111.09452 · 121 citations · Source: arxiv+semantic-scholarDespite great progress in object detection, most existing methods work only on a limited set of object categories, due to the tremendous human effort needed for bounding-box annotations of training data. To alleviate the problem, recent open vocabulary and zero-shot detection methods attempt to detect novel object categories beyond those seen during training. They achieve this goal by training on a pre-defined base categories to induce generalization to novel objects. However, their potential is still constrained by the small set of base categories available for training. To enlarge the set of base classes, we propose a method to automatically generate pseudo bounding-box annotations of diverse objects from large-scale image-caption pairs. Our method leverages the localization ability of pre-trained vision-language models to generate pseudo bounding-box labels and then directly uses them for training object detectors. Experimental results show that our method outperforms the state-of-the-art open vocabulary detector by 8% AP on COCO novel categories, by 6.3% AP on PASCAL VOC, by 2.3% AP on Objects365 and by 2.8% AP on LVIS. Code is available at https://github.com/salesforce/PB-OVD.
Ran Xu, Mingfei Gao, Chen Xing, Juan Carlos Niebles, Junnan Li, Wenhao Liu, Caiming Xiong · 7 authors totalAugmenting Deep Classifiers with Polynomial Neural Networks
European Conference on Computer Vision · DOI 10.1007/978-3-031-19806-9_40 · arXiv 2104.07916 · 32 citations · Source: semantic-scholar. Deep neural networks have been the driving force behind the success in classification tasks, e.g., object and audio recognition. Impressive results and generalization have been achieved by a variety of recently proposed architectures, the majority of which are seemingly disconnected. In this work, we cast the study of deep classifiers under a unifying framework. In particular, we express state-of-the-art architectures (e.g., residual and non-local networks) in the form of different degree polynomials of the input. Our framework provides insights on the inductive biases of each model and enables natural extensions building upon their polynomial nature. The efficacy of the proposed models is evaluated on standard image and audio classification benchmarks. The expressivity of the proposed models is highlighted both in terms of increased model performance as well as model compression. Lastly, the extensions allowed by this taxonomy showcase benefits in the presence of limited data and long-tailed data distributions. We expect this taxonomy to provide links between existing domain-specific architectures. The source code is available at https://github. com/grigorisg9gr/polynomials-for-augmenting-NNs . collection of state-of-the-art neural architectures as polynomials. Our unifying framework sheds light on the inductive bias of each
Jean Kossaifi, Grigorios G. Chrysos, Markos Georgopoulos, Jiankang Deng, Yannis Panagakis, Anima Anandkumar · 6 authors totalPractical Rule-Based Qualitative Temporal Reasoning for the Semantic Web
RuleML+RR · DOI 10.1007/978-3-030-91167-6_13 · 1 citations · Source: crossref+semantic-scholarRosario Uceda-Sosa, Guilherme Lima, Marcelo de Oliveira Costa Machado, Rosario A. Uceda-Sosa, M. Moreno · 5 authors totalThe Attention Economy and the Impact of Artificial Intelligence
DOI 10.1007/978-3-030-86144-5_18 · 23 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates, Usama M. Fayyad · 3 authors totalThreshold Schnorr with Stateless Deterministic Signing from Standard Assumptions
Advances in Cryptology — CRYPTO · DOI 10.1007/978-3-030-84242-0_6 · Source: springer+dblp+personal-first-partyFrançois Garillot, Francois Garillot, Yashvanth Kondi, Payman Mohassel, Valeria Nikolaenko · 5 authors totalNon-interactive Half-Aggregation of EdDSA and Variants of Schnorr Signatures
Topics in Cryptology — CT-RSA · DOI 10.1007/978-3-030-75539-3_24 · Source: springer+dblp+personal-first-partyFrançois Garillot, Konstantinos Chalkias, Francois Garillot, Yashvanth Kondi, Valeria Nikolaenko · 5 authors totalTraining and Evaluation of Word Embedding Models for Azerbaijani Language
Advances in Intelligent Systems and Computing (MIDI 2020) · DOI 10.1007/978-3-030-74728-2_4 · 9 citations · Source: crossref+openalex+dblpJavid Huseynov, Kamran Huseynov, Umid Suleymanov, Samir Rustamov, Javid J. Huseynov · 5 authors totalComplement Lexical Retrieval Model with Semantic Residual Embeddings
European Conference on Information Retrieval · DOI 10.1007/978-3-030-72113-8_10 · 118 citations · Source: semantic-scholarTongfei Chen, Luyu Gao, Zhuyun Dai, Zhen Fan, Benjamin Van Durme, Jamie Callan · 6 authors totalAI & Human Values
Lecture notes in computer science · DOI 10.1007/978-3-030-69128-8_6 · 29 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Laurence Devillers, Françoise Fogelman‐Soulié, Ricardo Baeza‐Yates · 4 authors totalBiomedical Named Entity Recognition at Scale
Lecture notes in computer science · DOI 10.1007/978-3-030-68763-2_48 · arXiv 2011.06315 · 4 citations · Source: openalexDavid Talby, Veysel Kocaman · 2 authors totalBlockchain Trading and Exchange
The Palgrave Handbook of Technological Finance · DOI 10.1007/978-3-030-65117-6_14 · 1 citations · Source: semantic-scholarCameron Pfiffer, Hugo E Benedetti, S. McKeon · 3 authors totalConvolutional Neural Networks with Swift for TensorFlow: Image Recognition and Dataset Categorization
Apress · DOI 10.1007/978-1-4842-6168-2 · Source: springer+author-first-partybrett koonce · 1 author totalChanging Health-Related Behaviors 6: Analysis, Interpretation, and Application of Big Data.
Methods in molecular biology · DOI 10.1007/978-1-0716-1138-8_34 · 0 citations · Source: semantic-scholarRandy Giffen, D. Bryant · 2 authors totalCategorizing metadata to help mobilize computable biomedical knowledge
Learning Health Systems · DOI 10.1002/lrh2.10271 · 29 citations · Source: openalex+first-party-career-authorityMark Samuel Tuttle, Brian S. Alper, Allen Flynn, Bruce E. Bray, Marisa Conte, Christina Eldredge, Sigfried Gold, Robert A. Greenes · 16 authors totalAutomatic analysis and validation of digitized speech markers in Lewy body spectrum diseases with Alzheimer’s disease co‐pathology
Alzheimer s & Dementia · DOI 10.1002/alz.053264 · 0 citations · Source: openalex+first-party-career-authorityMark Liberman, Sanjana Shellikeri, Sunghye Cho, Erica Howard, Yvonne Balganorth, Daniel Weintraub, Eddie B Lee, John Q. Trojanowski · 11 authors totalAutomatic classification of AD versus FTLD pathology using speech analysis in a biologically confirmed cohort
Alzheimer s & Dementia · DOI 10.1002/alz.052270 · 2 citations · Source: openalex+first-party-career-authorityMark Liberman, Sunghye Cho, Sanjana Shellikeri, Sharon Ash, Murray Grossman, Naomi Nevler · 7 authors totalDeeper Clinical Document Understanding Using Relation Extraction
arXiv · DOI 10.48550/arxiv.2112.13259 · arXiv 2112.13259 · 7 citations · Source: openalexThe surging amount of biomedical literature & digital clinical records presents a growing need for text mining techniques that can not only identify but also semantically relate entities in unstructured data. In this paper we propose a text mining framework comprising of Named Entity Recognition (NER) and Relation Extraction (RE) models, which expands on previous work in three main ways. First, we introduce two new RE model architectures -- an accuracy-optimized one based on BioBERT and a speed-optimized one utilizing crafted features over a Fully Connected Neural Network (FCNN). Second, we evaluate both models on public benchmark datasets and obtain new state-of-the-art F1 scores on the 2012 i2b2 Clinical Temporal Relations challenge (F1 of 73.6, +1.2% over the previous SOTA), the 2010 i2b2 Clinical Relations challenge (F1 of 69.1, +1.2%), the 2019 Phenotype-Gene Relations dataset (F1 of 87.9, +8.5%), the 2012 Adverse Drug Events Drug-Reaction dataset (F1 of 90.0, +6.3%), and the 2018 n2c2 Posology Relations dataset (F1 of 96.7, +0.6%). Third, we show two practical applications of this framework -- for building a biomedical knowledge graph and for improving the accuracy of mapping entities to clinical codes. The system is built using the Spark NLP library which provides a production-grade, natively scalable, hardware-optimized, trainable & tunable NLP framework.
David Talby, Hasham Ul Haq, Veysel Kocaman · 3 authors totalScaling Language Models: Methods, Analysis & Insights from Training Gopher
arXiv · arXiv 2112.11446 · 1,655 citations · Source: arxiv+semantic-scholarLanguage modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world. In this paper, we present an analysis of Transformer-based language model performance across a wide range of model scales -- from models with tens of millions of parameters up to a 280 billion parameter model called Gopher. These models are evaluated on 152 diverse tasks, achieving state-of-the-art performance across the majority. Gains from scale are largest in areas such as reading comprehension, fact-checking, and the identification of toxic language, but logical and mathematical reasoning see less benefit. We provide a holistic analysis of the training dataset and model's behaviour, covering the intersection of model scale with bias and toxicity. Finally we discuss the application of language models to AI safety and the mitigation of downstream harms.
Erich Elsen, Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides · 80 authors totalMore Reviews May Not Help: Evidence from Incentivized First Reviews on Airbnb
arXiv · arXiv 2112.09783 · Source: arxiv+author-first-partyDave Holtz, Andrey Fradkin, David Holtz · 3 authors totalFLoRA: Single-shot Hyper-parameter Optimization for Federated Learning
arXiv.org · arXiv 2112.08524 · 28 citations · Source: semantic-scholar+arxivWe address the relatively unexplored problem of hyper-parameter optimization (HPO) for federated learning (FL-HPO). We introduce Federated Loss suRface Aggregation (FLoRA), the first FL-HPO solution framework that can address use cases of tabular data and gradient boosting training algorithms in addition to stochastic gradient descent/neural networks commonly addressed in the FL literature. The framework enables single-shot FL-HPO, by first identifying a good set of hyper-parameters that are used in a **single** FL training. Thus, it enables FL-HPO solutions with minimal additional communication overhead compared to FL training without HPO. Our empirical evaluation of FLoRA for Gradient Boosted Decision Trees on seven OpenML data sets demonstrates significant model accuracy improvements over the considered baseline, and robustness to increasing number of parties involved in FL-HPO training.
Nathalie Baracaldo, Yi Zhou, Parikshit Ram, Theodoros Salonidis, Horst Samulowitz, Heiko Ludwig · 6 authors totalGLaM: Efficient Scaling of Language Models with Mixture-of-Experts
CoRR · arXiv 2112.06905 · Source: first-party+openalexMaarten Bosma, Nan Du, Yanping Huang, Andrew M. Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun · 27 authors totalStep-unrolled Denoising Autoencoders for Text Generation
ICLR · arXiv 2112.06749 · 153 citations · Source: arxiv+semantic-scholarIn this paper we propose a new generative model of text, Step-unrolled Denoising Autoencoder (SUNDAE), that does not rely on autoregressive models. Similarly to denoising diffusion techniques, SUNDAE is repeatedly applied on a sequence of tokens, starting from random inputs and improving them each time until convergence. We present a simple new improvement operator that converges in fewer iterations than diffusion methods, while qualitatively producing better samples on natural language datasets. SUNDAE achieves state-of-the-art results (among non-autoregressive methods) on the WMT'14 English-to-German translation task and good qualitative results on unconditional language modeling on the Colossal Cleaned Common Crawl dataset and a dataset of Python code from GitHub. The non-autoregressive nature of SUNDAE opens up possibilities beyond left-to-right prompted generation, by filling in arbitrary blank patterns in a template.
Erich Elsen, Nikolay Savinov, Junyoung Chung, Mikolaj Binkowski, Aaron van den Oord · 5 authors total