Papers.
Research connected to its authors, projects, companies, talks, events, and the rest of the graph.
Add a paper ↗(De/Re)-Compositions Expressed Systematically via MDH-Based Schedules
DOI 10.1145/3578360.3580269 · 3 citations · Source: openalexWe introduce a new scheduling language, based on the formalism of Multi-Dimensional Homomorphisms (MDH). In contrast to existing scheduling languages, our MDH-based language is designed to systematically "de-compose" computations for the memory and core hierarchies of architectures, and "re-compose" the computed intermediate results back to the final result -- we say "(de/re)-composition" for short. We argue that our scheduling langauge is easy to use and yet expressive enough to express well-performing (de/re)-compositions of popular related approaches, e.g., the TVM compiler, for MDH-supported computations (such as linear algebra routines and stencil computations). Moreover, our language is designed as auto-tunable, i.e., any optimization decision can optionally be left to the auto-tuning engine of our system, and our system can automatically recommend schedules for the user, based on its auto-tuning capabilities. Also, by relying on the MDH approach, we can formally guarantee the correctness of optimizations expressed in our language, thereby further enhancing user experience. Our experiments on GPU and CPU confirm that we can express optimizations that cannot be expressed straightforwardly (or at all) in TVM's scheduling language, thereby achieving higher performance than TVM, and also vendor libraries provided by NVIDIA and Intel, for time-intensive computations used in real-world deep learning neural networks.
Denys Shabalin, Ari Rasch, Richard Schulze, Anne C. Elster, Sergei Gorlatch, Mary Hall · 6 authors totalFor-Each Operations in Collaborative Apps
PaPoC at EuroSys · DOI 10.1145/3578358.3591323 · Source: acm+orcid+dblp+cmu-career-authorityHeather, Matthew Weidner, Ria Pradeep, Benito Geordie, Heather Miller · 5 authors totalDemonstration of Geyser: Provenance Extraction and Applications over Data Science Scripts
SIGMOD (demo) · DOI 10.1145/3555041.3589717 · 4 citations · Source: semantic-scholarAshvin Agrawal, Fotis Psallidas, Megan Leszczynski, Mohammad Hossein Namaki, Avrilia Floratou, Konstantinos Karanasos, Subru Krishnan, Pavle Subotic · 11 authors totalBJC Sparks: A New Functional-First Middle School CS Curriculum
54th ACM Technical Symposium on Computer Science Education · DOI 10.1145/3545945.3569842 · Source: acm+author-first-partyDave Briccetti, Dan Garcia, Mary Fries, Michael Ball, Pamela Fox, Deanna Gelosi, Lauren Mock, Della Dastur · 9 authors totalHuman-Centered Responsible Artificial Intelligence: Current & Future Trends
DOI 10.1145/3544549.3583178 · 44 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Mohammad Tahaei, Marios Constantinides, Daniele Quercia, Seán Kennedy, Michael Müller, Simone Stumpf, Q. Vera Liao · 16 authors totalSpeakFaster Observer: Long-Term Instrumentation of Eye-Gaze Typing for Measuring AAC Communication.
CHI Extended Abstracts · DOI 10.1145/3544549.3573870 · Source: dblpKatrin Tomanek, Shanqing Cai, Subhashini Venugopalan, Shaun K. Kane, Meredith Ringel Morris, Richard Cave, Robert L. MacDonald, Jon Campbell · 12 authors total13th Temporal Web Analytics Workshop (TempWeb) Overview
DOI 10.1145/3543873.3589742 · 1 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Marc Spaniol, Ricardo Baeza‐Yates, Omar Alonso · 4 authors totalIntent-Aware Propensity Estimation via Click Pattern Stratification
WWW Companion · DOI 10.1145/3543873.3587610 · 0 citations · Source: openalex+orcidCounterfactual learning to rank via inverse propensity weighting is the most popular approach to train ranking models using biased implicit user feedback from logged search data. Standard click propensity estimation techniques rely on simple models of user browsing behavior that primarily account for the attributes of the presentation context that affect whether the relevance of an item to the search context is observed. Most notably, the inherent effect of the listwise presentation of the items on users’ propensity for engagement is captured in the position of the presented items on the search result page. In this work, we enrich this position bias based click propensity model by proposing an observation model that further incorporates the underlying search intent, as reflected in the user’s click pattern in the search context. Our approach does not require an intent prediction model based on the content of the search context. Instead, we rely on a simple, yet effective, non-causal estimate of the user’s browsing intent from the number of click events in the search context. We empirically characterize the distinct rank decay patterns of the estimated click propensities in the characterized intent classes. In particular, we demonstrate a sharper decay of click propensities in top ranks for the intent class identified by sparse user clicks and the higher likelihood of observing clicks in lower ranks for the intent class identified by higher number of user clicks. We show that the proposed intent-aware propensity estimation technique helps with training ranking models with more effective personalization and generalization power through empirical results for a ranking task in a major e-commerce platform.
Alex Cozzi, Ehsan Ebrahimzadeh, Abraham Bagherjeiran · 3 authors totalSoftware Engineering of Machine Learning Systems
Communications of the ACM · DOI 10.1145/3539783 · 12 citations · Source: openalexSeeking to make machine learning more dependable.
Peter Norvig, Charles L. Isbell, Michael L. Littman · 3 authors totalSearching for Reliable Facts over a Medical Knowledge Base
SIGIR · DOI 10.1145/3539618.3591822 · 11 citations · Source: dblp+semantic-scholarThis work presents CoreKB, a Web platform for searching reliable facts over gene expression-cancer associations Knowledge Base (KB). It provides search capabilities over an RDF graph using natural language queries, structured facets, and autocomplete. CoreKB is designed to be intuitive and easy to use for healthcare professionals, medical researchers, and clinicians. The system offers the user a comprehensive overview of the scientific evidence supporting a medical fact. It provides a quantitative comparison between the possible gene-cancer associations a particular fact can reflect.
Omar Alonso, Fabio Giachelle, Stefano Marchesin, G. Silvello · 4 authors totalThe effect of phrasing, speech rate, and information structure on tonal coarticulation in spontaneous Cantonese
The Journal of the Acoustical Society of America · DOI 10.1121/10.0018896 · 0 citations · Source: openalex+first-party-career-authorityMark Liberman, Xin Gao · 2 authors totalExploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks
2024 IEEE Security and Privacy Workshops (SPW) · DOI 10.1109/SPW63631.2024.00018 · arXiv 2302.05733 · 393 citations · Source: semantic-scholarRecent advances in instruction-following large language models (LLMs) have led to dramatic improvements in a range of NLP tasks. Unfortunately, we find that the same improved capabilities amplify the dual-use risks for malicious purposes of these models. Dual-use is difficult to prevent as instruction-following capabilities now enable standard attacks from computer security. The capabilities of these instruction-following LLMs provide strong economic incentives for dual-use by malicious actors. In particular, we show that instruction-following LLMs can produce targeted malicious content, including hate speech and scams, bypassing in-the-wild defenses implemented by LLM API vendors. Our analysis shows that this content can be generated economically and at cost of $125-500 \times$ cheaper than human effort alone. Together, our findings suggest that LLMs will increasingly attract more sophisticated adversaries and attacks, and addressing these attacks may require new approaches to mitigations.
Carlos Guestrin, Daniel Kang, Xuechen Li, Ion Stoica, M. Zaharia, Tatsunori Hashimoto · 6 authors totalA Deployment-First Methodology to Mechanism Design and Refinement in Distributed Systems
DOI 10.1109/percomworkshops56833.2023.10150355 · 1 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Martijn de Vos, Georgy Ishmaev, Stefanie Roos · 4 authors totalSustainable Cooperation in Peer-To-Peer Networks
DOI 10.1109/lcn58197.2023.10223360 · 2 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Bulat Nasrulin, Rowdy Chotkan · 3 authors totalMultiyear Mapping of Water Demand at Crop Level: An End-to-End Workflow Based on High-Resolution Crop Type Maps and Meteorological Data
IEEE J. Sel. Top. Appl. Earth Obs. Remote. Sens. · DOI 10.1109/JSTARS.2023.3294107 · Source: dblp+first-party-career-authorityJim Dowling, Giulio Weikmann, Daniele Marinelli, Claudia Paris, Silke Migdall, Eva Gleisberg, Florian Appel, Heike Bach · 9 authors totalEnhancing Human Cobot Interaction using Natural Language Processing
2023 IEEE 4th International Multidisciplinary Conference on Engineering Technology (IMCET) · DOI 10.1109/IMCET59736.2023.10368263 · 5 citations · Source: semantic-scholarThis paper presents a unique approach to enhance Human-Cobot Interaction using Natural Language Processing. We integrate the CoboVox voice recognition system into the UR3e cobot interface, enabling voice-activated engagement with collaborative robots. Our methodology includes vocabulary building, language processing model analysis, and practical implementation, empowering non-expert robot operators to efficiently program and control cobots through voice commands. This innovation holds potential benefits for Industry 4.0 by promoting responsible implementation and improved human-robot collaboration.
Gautam Siwach, Cheryl Li · 2 authors totalDeformer: Dynamic Fusion Transformer for Robust Hand Pose Estimation
IEEE International Conference on Computer Vision · DOI 10.1109/ICCV51070.2023.02157 · arXiv 2303.04991 · 33 citations · Source: arxiv+semantic-scholarAccurately estimating 3D hand pose is crucial for understanding how humans interact with the world. Despite remarkable progress, existing methods often struggle to generate plausible hand poses when the hand is heavily occluded or blurred. In videos, the movements of the hand allow us to observe various parts of the hand that may be occluded or blurred in a single frame. To adaptively leverage the visual clue before and after the occlusion or blurring for robust hand pose estimation, we propose the Deformer: a framework that implicitly reasons about the relationship between hand parts within the same image (spatial dimension) and different timesteps (temporal dimension). We show that a naive application of the transformer self-attention mechanism is not sufficient because motion blur or occlusions in certain frames can lead to heavily distorted hand features and generate imprecise keys and queries. To address this challenge, we incorporate a Dynamic Fusion Module into Deformer, which predicts the deformation of the hand and warps the hand mesh predictions from nearby frames to explicitly support the current frame estimation. Furthermore, we have observed that errors are unevenly distributed across different hand parts, with vertices around fingertips having disproportionately higher errors than those around the palm. We mitigate this issue by introducing a new loss function called maxMSE that automatically adjusts the weight of every vertex to focus the model on critical hand parts. Extensive experiments show that our method significantly outperforms state-of-the-art methods by 10%, and is more robust to occlusions (over 14%).
Ran Xu, Qichen Fu, Xingyu Liu, Juan Carlos Niebles, Kris Kitani · 5 authors totalGlueGen: Plug and Play Multi-modal Encoders for X-to-image Generation
IEEE International Conference on Computer Vision · DOI 10.1109/ICCV51070.2023.02110 · arXiv 2303.10056 · 31 citations · Source: arxiv+semantic-scholarText-to-image (T2I) models based on diffusion processes have achieved remarkable success in controllable image generation using user-provided captions. However, the tight coupling between the current text encoder and image decoder in T2I models makes it challenging to replace or upgrade. Such changes often require massive fine-tuning or even training from scratch with the prohibitive expense. To address this problem, we propose GlueGen, which applies a newly proposed GlueNet model to align features from single-modal or multi-modal encoders with the latent space of an existing T2I model. The approach introduces a new training objective that leverages parallel corpora to align the representation spaces of different encoders. Empirical results show that GlueNet can be trained efficiently and enables various capabilities beyond previous state-of-the-art models: 1) multilingual language models such as XLM-Roberta can be aligned with existing T2I models, allowing for the generation of high-quality images from captions beyond English; 2) GlueNet can align multi-modal encoders such as AudioCLIP with the Stable Diffusion model, enabling sound-to-image generation; 3) it can also upgrade the current text encoder of the latent diffusion model for challenging case generation. By the alignment of various feature representations, the GlueNet allows for flexible and efficient integration of new functionality into existing T2I models and sheds light on X-to-image (X2I) generation.1
Ran Xu, Can Qin, Ning Yu, Chen Xing, Shu Zhang, Zeyuan Chen, S. Ermon, Yun Fu · 9 authors totalMake-An-Animation: Large-Scale Text-conditional 3D Human Motion Generation
ICCV 2023 · DOI 10.1109/iccv51070.2023.01381 · arXiv 2305.09662 · 74 citations · Source: semantic-scholarText-guided human motion generation has drawn significant interest because of its impactful applications spanning animation and robotics. Recently, application of diffusion models for motion generation has enabled improvements in the quality of generated motions. However, existing approaches are limited by their reliance on relatively small-scale motion capture data, leading to poor performance on more diverse, in-the-wild prompts. In this paper, we introduce Make-An-Animation, a text-conditioned human motion generation model which learns more diverse poses and prompts from large-scale image-text datasets, enabling significant improvement in performance over prior works. Make-An-Animation is trained in two stages. First, we train on a curated large-scale dataset of (text, static pseudo-pose) pairs extracted from image-text datasets. Second, we fine-tune on motion capture data, adding additional layers to model the temporal dimension. Unlike prior diffusion models for motion generation, Make-An-Animation uses a U-Net architecture similar to recent text-to-video generation models. Human evaluation of motion realism and alignment with input text shows that our model reaches state-of-the-art performance on text-to-motion generation. Generated samples can be viewed at https://azadis.github.io/make-an-animation.
Sonal Gupta, S. Azadi, Akbar Shah, Thomas Hayes, Devi Parikh · 5 authors totalAn Analysis of Degenerating Speech Due to Progressive Dysarthria on ASR Performance.
ICASSP · DOI 10.1109/icassp49357.2023.10097195 · Source: dblpKatrin Tomanek, Katie Seaver, Pan-Pan Jiang, Richard Cave, Lauren Harrell, Jordan R. Green · 6 authors totalCompiler-Supported Selective Software Fault Tolerance
IEEE Conference on Dependable and Secure Computing · DOI 10.1109/DSC61021.2023.10354221 · Source: ieee+bilkent-repositoryHakan Tekgul, Tuncer Turhan, Ozcan Ozturk · 3 authors totalULIP-2: Towards Scalable Multimodal Pre-Training for 3D Understanding
Computer Vision and Pattern Recognition · DOI 10.1109/CVPR52733.2024.02558 · arXiv 2305.08275 · 261 citations · Source: arxiv+semantic-scholarRecent advancements in multimodal pretraining have shown promising efficacy in 3D representation learning by aligning multimodal features across 3D shapes, their 2D counterparts, and language descriptions. However, the methods used by existing frameworks to curate such multimodal data, in particular language descriptions for 3D shapes, are not scalable, and the collected language descriptions are not diverse. To address this, we introduce ULIP-2, a simple yet effective tri-modal pretraining framework that leverages large multimodal models to automatically generate holistic language descriptions for 3D shapes. It only needs 3D data as input, eliminating the need for any manual 3D annotations, and is therefore scalable to large datasets. ULIP-2 is also equipped with scaled-up backbones for better multimodal representation learning. We conduct experiments on two large-scale 3D datasets, Objaverse and ShapeNet, and augment them with tri-modal datasets of 3D point clouds, images, and language for training ULIP-2. Experiments show that ULIP-2 demonstrates substantial benefits in three downstream tasks: zero-shot 3D classification, standard 3D classification with fine-tuning, and 3D captioning (3D-to-language generation). It achieves a new SOTA of 50.6% (top-1) on Objaverse-LVIS and 84.7% (top-1) on ModelNet40 in zero-shot classification. In the ScanObjectNN benchmark for standard fine-tuning, ULIP-2 reaches an overall accuracy of 91.5% with a compact model of only 1.4 million parameters. ULIP-2 sheds light on a new paradigm for scalable multimodal 3D representation learning without human annotations and shows significant improvements over existing baselines. The code and datasets are released at https://github.com/salesforce/ULIP.
Ran Xu, Le Xue, Ning Yu, Shu Zhang, Junnan Li, Roberto Mart'in-Mart'in, Jiajun Wu, Caiming Xiong · 10 authors totalHIVE: Harnessing Human Feedback for Instructional Visual Editing
Computer Vision and Pattern Recognition · DOI 10.1109/CVPR52733.2024.00862 · arXiv 2303.09618 · 197 citations · Source: arxiv+semantic-scholarIncorporating human feedback has been shown to be crucial to align text generated by large language models to human preferences. We hypothesize that state-of-the-art instructional image editing models, where outputs are generated based on an input image and an editing instruction, could similarly benefit from human feedback, as their outputs may not adhere to the correct instructions and preferences of users. In this paper, we present a novel framework to harness human feedback for instructional visual editing (HIVE). Specifically, we collect human feedback on the edited images and learn a reward function to capture the underlying user preferences. We then introduce scalable diffusion model fine-tuning methods that can incorporate human preferences based on the estimated reward. Besides, to mitigate the bias brought by the limitation of data, we contribute a new 1.1M training dataset, a 3.6K reward dataset for rewards learning, and a 1 K evaluation dataset to boost the performance of instructional image editing. We conduct extensive empirical experiments quantitatively and qualitatively, showing that HIVE is favored over previous state-of-the-art instructional image editing approaches by a large margin.
Ran Xu, Shu Zhang, Xinyi Yang, Yihao Feng, Can Qin, Chia-Chih Chen, Ning Yu, Zeyuan Chen · 12 authors totalMask-Free OVIS: Open-Vocabulary Instance Segmentation without Manual Mask Annotations
Computer Vision and Pattern Recognition · DOI 10.1109/CVPR52729.2023.02254 · arXiv 2303.16891 · 25 citations · Source: arxiv+semantic-scholarExisting instance segmentation models learn task-specific information using manual mask annotations from base (training) categories. These mask annotations require tremendous human effort, limiting the scalability to annotate novel (new) categories. To alleviate this problem, Open-Vocabulary (OV) methods leverage large-scale image-caption pairs and vision-language models to learn novel categories. In summary, an OV method learns task-specific information using strong supervision from base annotations and novel category information using weak supervision from image-captions pairs. This difference between strong and weak supervision leads to overfitting on base categories, resulting in poor generalization towards novel categories. In this work, we overcome this issue by learning both base and novel categories from pseudo-mask annotations generated by the vision-language model in a weakly supervised manner using our proposed Mask-free OVIS pipeline. Our method automatically generates pseudo-mask annotations by leveraging the localization ability of a pre-trained vision-language model for objects present in image-caption pairs. The generated pseudo-mask annotations are then used to supervise an instance segmentation model, freeing the entire pipeline from any labour-expensive instance-level annotations and overfitting. Our extensive experiments show that our method trained with just pseudo-masks significantly improves the mAP scores on the MS-COCO dataset and OpenImages dataset compared to the recent state-of-the-art methods trained with manual masks. Codes and models are provided in https://vibashan.github.io/ovis-web/.
Ran Xu, V. Vibashan, Ning Yu, Chen Xing, Can Qin, Mingfei Gao, Juan Carlos Niebles, Vishal M. Patel · 8 authors totalMasked Autoencoding Does Not Help Natural Language Supervision at Scale
Computer Vision and Pattern Recognition · DOI 10.1109/CVPR52729.2023.02244 · arXiv 2301.07836 · 8 citations · Source: semantic-scholarSelf supervision and natural language supervision have emerged as two exciting ways to train general purpose image encoders which excel at a variety of downstream tasks. Recent works such as M3AE [31] and SLIP [63] have suggested that these approaches can be effectively combined, but most notably their results use small <20M examples) pre-training datasets and don't effectively reflect the large-scale regime (> 100M samples) that is commonly used for these approaches. Here we investigate whether a similar approach can be effective when trained with a much larger amount of data. We find that a combination of two state of the art approaches: masked autoencoders, MAE [37] and contrastive language image pretraining, CLIP [68] provides a benefit over CLIP when trained on a corpus of 11.3M image-text pairs, but little to no benefit (as evaluated on a suite of common vision tasks) over CLIP when trained on a large corpus of 1.4B images. Our work provides some much needed clarity into the effectiveness (or lack thereof) of self supervision for large-scale image-text training.
Vaishaal Shankar, Floris Weers, Angelos Katharopoulos, Yinfei Yang, Tom Gunter · 5 authors totalGame-map Pathfinding with Per-Problem Selection of Synthesized Heuristics.
CoG · DOI 10.1109/cog57401.2023.10333175 · Source: dblp+ubc-authorityRamon Lawrence, Vadim Bulitko · 2 authors totalIDMU: Impact Driven Machine Unlearning
BigData Congress [Services Society] · DOI 10.1109/BIGDATA59044.2023.10386841 · 2 citations · Source: crossref+semantic-scholarEnterprise organizations have large amounts of data which is utilized by multiple Machine Learning (ML) models over various software frameworks. These models provide trends and insights from the data that can help enterprises define business rules around their processes. However, if certain aspects of this data are removed from the datasets, it could influence the business rules and policies in place. When a user requests data to be removed, the model retraining may be required called Machine Unlearning (MU). Recent research works in the area of MU include different methods of retraining the machine learning models. It turns out that there is lack of work in removing certain aspects of data, and quantifying its impact on the models. This paper aspires to provide a novel methodology IDMU (Impact Driven Machine Unlearning) that performs quantification of the impact of data removal requests while performing MU. Our method provides recommendations for data removal requests, factoring in underlying features of data. The results from the industrial application and evaluation of our method on a financial services dataset are encouraging. The overall IDMU had a mean MAPE of 10.25% over a set of 120 data removal requests. It also saved ~1900 hours of model retraining time by factoring in urgency and impact of data removal requests over a period of three years.
Ruchi Mahindru, Shubhi Asthana, Bing Zhang, R. Mahindru, I. Banipal, Pawan Chowdhary · 6 authors totalMeasuring Bias
DOI 10.1109/bigdata59044.2023.10386679 · 1 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Aida Sharif Rohani, Ricardo Baeza‐Yates · 3 authors totalDigital markers of motor speech impairments in natural speech of patients with ALS-FTD spectrum disorders
medRxiv · DOI 10.1101/2023.04.29.23289308 · 0 citations · Source: openalex+first-party-career-authorityMark Liberman, Sanjana Shellikeri, Sunghye Cho, Sharon Ash, Carmen Gonzalez-Recober, Corey T. McMillan, Lauren Elman, Colin Quinn · 15 authors totalRNAget: an API to securely retrieve RNA quantifications
Bioinformatics · DOI 10.1093/bioinformatics/btad126 · 5 citations · Source: semantic-scholarAbstract Summary Large-scale sharing of genomic quantification data requires standardized access interfaces. In this Global Alliance for Genomics and Health project, we developed RNAget, an API for secure access to genomic quantification data in matrix form. RNAget provides for slicing matrices to extract desired subsets of data and is applicable to all expression matrix-format data, including RNA sequencing and microarrays. Further, it generalizes to quantification matrices of other sequence-based genomics such as ATAC-seq and ChIP-seq. Availability and implementation https://ga4gh-rnaseq.github.io/schema/docs/index.html.
Alyssa Morrow, Sean Upchurch, Emilio Palumbo, Jeremy Adams, David Bujold, Guillaume Bourque, Jared L. Nedzel, Keenan Graham · 68 authors totalDigital markers of motor speech impairments in spontaneous speech of patients with ALS-FTD spectrum disorders
Amyotrophic Lateral Sclerosis and Frontotemporal Degeneration · DOI 10.1080/21678421.2023.2288106 · 2 citations · Source: openalex+first-party-career-authorityMark Liberman, Sanjana Shellikeri, Sunghye Cho, Sharon Ash, Carmen Gonzalez-Recober, Corey T. McMillan, Lauren Elman, Colin Quinn · 15 authors totalFederated benchmarking of medical artificial intelligence with MedPerf
Nature Machine Intelligence · DOI 10.1038/s42256-023-00652-2 · 161 citations · Source: openalexMedical artificial intelligence (AI) has tremendous potential to advance healthcare by supporting and contributing to the evidence-based practice of medicine, personalizing patient treatment, reducing costs, and improving both healthcare provider and patient experience. Unlocking this potential requires systematic, quantitative evaluation of the performance of medical AI models on large-scale, heterogeneous data capturing diverse patient populations. Here, to meet this need, we introduce MedPerf, an open platform for benchmarking AI models in the medical domain. MedPerf focuses on enabling federated evaluation of AI models, by securely distributing them to different facilities, such as healthcare organizations. This process of bringing the model to the data empowers each facility to assess and verify the performance of AI models in an efficient and human-supervised process, while prioritizing privacy. We describe the current challenges healthcare and AI communities face, the need for an open platform, the design philosophy of MedPerf, its current implementation status and real-world deployment, our roadmap and, importantly, the use of MedPerf with multiple international institutions within cloud-based technology and on-premises scenarios. Finally, we welcome new contributions by researchers and organizations to further strengthen MedPerf as an open benchmarking platform.
David Talby, Alexandros Karargyris, Renato Umeton, Micah Sheller, Alejandro Aristizábal, Johnu George, Anna Wuest, Sarthak Pati · 74 authors totalNeural operators for accelerating scientific simulations and design
Nature Reviews Physics · DOI 10.1038/s42254-024-00712-5 · arXiv 2309.15325 · 445 citations · Source: semantic-scholarScientific discovery and engineering design are currently limited by the time and cost of physical experiments. Numerical simulations are an alternative approach but are usually intractable for complex real-world problems. Artificial intelligence promises a solution through fast data-driven surrogate models. In particular, neural operators present a principled framework for learning mappings between functions defined on continuous domains, such as spatiotemporal processes and partial differential equations. Neural operators can extrapolate and predict solutions at new locations unseen during training. They can be integrated with physics and other domain constraints enforced at finer resolutions to obtain high-fidelity solutions and good generalization. Neural operators are differentiable, so they can directly optimize parameters for inverse design and other inverse problems. Neural operators can therefore augment, or even replace, existing numerical simulators in many applications, such as computational fluid dynamics, weather forecasting and material modelling, providing speedups of four to five orders of magnitude. Neural operators learn mappings between functions on continuous domains, such as spatiotemporal processes and partial differential equations, offering a fast, data-driven surrogate model solution for otherwise intractable numerical simulations of complex real-world problems.
Jean Kossaifi, Kamyar Azzizadenesheli, Nikola B. Kovachki, Zong-Yi Li, Miguel Liu-Schiaffini, Anima Anandkumar · 6 authors totalLarge language models generate functional protein sequences across diverse families
Nature Biotechnology · DOI 10.1038/s41587-022-01618-2 · 1,131 citations · Source: semantic-scholarA generative deep-learning model designs artificial proteins with desired enzymatic activities. Deep-learning language models have shown promise in various biotechnological applications, including protein design and engineering. Here we describe ProGen, a language model that can generate protein sequences with a predictable function across large protein families, akin to generating grammatically and semantically correct natural language sentences on diverse topics. The model was trained on 280 million protein sequences from >19,000 families and is augmented with control tags specifying protein properties. ProGen can be further fine-tuned to curated sequences and tags to improve controllable generation performance of proteins from families with sufficient homologous samples. Artificial proteins fine-tuned to five distinct lysozyme families showed similar catalytic efficiencies as natural lysozymes, with sequence identity to natural proteins as low as 31.4%. ProGen is readily adapted to diverse protein families, as we demonstrate with chorismate mutase and malate dehydrogenase.
Richard Socher, Ali Madani, Ben Krause, E. Greene, Subu Subramanian, Benjamin P. Mohr, J. Holton, J. L. Olmos · 12 authors totalLECTURE HELD AT THE ACADEMIA EUROPAEA BUILDING BRIDGES CONFERENCE 2022
European Review · DOI 10.1017/s1062798723000145 · 9 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates · 2 authors totalydata-profiling: Accelerating data-centric AI with high-quality data
Neurocomputing · DOI 10.1016/j.neucom.2023.126585 · 25 citations · Source: dblpFabiana Clemente, Gonçalo Martins Ribeiro, Alexandre Quemy, Miriam Seoane Santos, Ricardo Cardoso Pereira, Alex Barros · 6 authors totalDeScan: Censorship-resistant indexing and search for Web3
Future Generation Computer Systems · DOI 10.1016/j.future.2023.11.008 · 3 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Martijn de Vos, Georgy Ishmaev · 3 authors totalWeb3 Sybil avoidance using network latency
Computer Networks · DOI 10.1016/j.comnet.2023.109701 · 8 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Quinten Stokkink, Can Umut Ileri, Dick Epema · 4 authors totalLane detection and path prediction in autonomous vehicle using deep learning
Intelligent Edge Computing for Cyber Physical Applications (Elsevier) · DOI 10.1016/B978-0-323-99412-5.00012-5 · 12 citations · Source: crossref+openalexJay Rodge, Renu Kachhoria, Swati Jaiswal, Meghana Lokhande · 4 authors totalMachine understanding and deep learning representation
Synthese · DOI 10.1007/s11229-022-03999-y · 15 citations · Source: semantic-scholarPractical ability manifested through robust and reliable task performance, as well as information relevance and well-structured representation, are key factors indicative of understanding in the philosophical literature. We explore these factors in the context of deep learning, identifying prominent patterns in how the results of these algorithms represent information. While the estimation applications of modern neural networks do not qualify as the mental activity of persons, we argue that coupling analyses from philosophical accounts with the empirical and theoretical basis for identifying these factors in deep learning representations provides a framework for discussing and critically evaluating potential machine understanding given the continually improving task performance enabled by such algorithms.
Mike Tamir, Michael Tamir, Elay Shech · 3 authors total