Papers.
Research connected to its authors, projects, companies, talks, events, and the rest of the graph.
Add a paper ↗Training Compute-Optimal Large Language Models
Advances in Neural Information Processing Systems 35 · DOI 10.52202/068431-2176 · arXiv 2203.15556 · 3,642 citations · Source: arxiv+semantic-scholarWe investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant. By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled. We test this hypothesis by training a predicted compute-optimal model, Chinchilla, that uses the same compute budget as Gopher but with 70B parameters and 4$\times$ more more data. Chinchilla uniformly and significantly outperforms Gopher (280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B) on a large range of downstream evaluation tasks. This also means that Chinchilla uses substantially less compute for fine-tuning and inference, greatly facilitating downstream usage. As a highlight, Chinchilla reaches a state-of-the-art average accuracy of 67.5% on the MMLU benchmark, greater than a 7% improvement over Gopher.
Erich Elsen, Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas · 22 authors totalType-Preserving Compilation of Class-Based Languages
EPFL PhD thesis · DOI 10.5075/epfl-thesis-10108 · arXiv 2307.05557 · Source: epfl+arxiv+personal-first-partyGuillaume Martres · 1 author totalExpecting Too Much from Our Machine Learning Models
Scientific Understanding and Representation (Routledge) · DOI 10.4324/9781003202905-30 · Source: crossrefMike Tamir, Elay Shech, Michael Tamir · 3 authors totalUnderstanding from Deep Learning Models in Context
Scientific Understanding and Representation (Routledge) · DOI 10.4324/9781003202905-28 · Source: crossrefMike Tamir, Michael Tamir, Elay Shech · 3 authors totalRetention of epigenetic and transcriptional memory among reprogrammed T cells (T-iPSCs) imparts selective advantage for T cell re-differentiation
Journal of Immunology · DOI 10.4049/jimmunol.208.supp.107.12 · 0 citations · Source: semantic-scholarVarious somatic cell types reprogramed into induced pluripotent stem cells (iPSCs) can be differentiated into T cells. It remains unclear what impact the input cells reprogramed into iPSCs have on T cell differentiation. We hypothesized that preservation of T cell genetic and/or epigenetic memory would bias iPSC clones toward T cell differentiation. CD8 naïve T cells (TN), CD8 stem memory T cells (TSM), and fibroblasts (FB) were reprogramed into iPSCs using nonintegrating Sendai virus to deliver Yamanaka factors. Input cells and iPSC clones were subjected to RNA-seq and ATAC-seq to assess whether epigenetic and transcriptomic memory are retained after reprogramming. Pluripotent associated genes (NANOG, LIN28A/B, SOX2) had increased expression in both T cell− and FB-iPSCs compared to their cells of origin. Of interest, gene enrichment analysis of fibroblasts and FB-iPSCs identified fibroblast related annotations, including migration and growth factors, in reprogramed iPSCs. Among CD8 TN− and TSM-iPSCs as a group, expression of only 41 genes and 6 genomic regions were retained, but these included known mature T cell genes, CD3ɛ, ZBTB7B, and CD95 as well as a region crucial for T cell signaling. Therefore, although much of the transcriptome and accessible genome is overridden in TN/TSM-iPSCs, there exists maintenance of both the transcriptome and chromatin accessibility after reprogramming. CD8 T-iPSCs were superior to FB-iPSCs in producing T progenitors in vitro. Highlighted here is the importance of starting populations for reprogramming and differentiation into T lineage cells. Leveraging the benefits of epigenetic and transcriptomic memory of input T cells could provide a more robust source of T cells for adoptive cell therapy. Supported in part by funds from NIH P01 CA065493 and R37 AI34495 from the National Cancer Institute and National Institute of Allergy and Infectious disease, respectively, the Childrens' Cancer Research Fund, and Chan-Zuckerberg Biohub.
Alyssa Morrow, Robin Lesley Williams, Jeremy Allred, Jakub Tolar, W. Nicholas Haining, Nir Yosef, Bruce R. Blazar · 7 authors totalFirst Sagittarius A* Event Horizon Telescope Results. VI. Testing the Black Hole Metric
The Astrophysical Journal Letters · DOI 10.3847/2041-8213/ac6756 · Source: iop+orcid+smithsonianGreg Lindahl, Event Horizon Telescope Collaboration · 2 authors totalFirst Sagittarius A* Event Horizon Telescope Results. IV. Variability, Morphology, and Black Hole Mass
The Astrophysical Journal Letters · DOI 10.3847/2041-8213/ac6736 · Source: iop+orcid+smithsonianGreg Lindahl, Event Horizon Telescope Collaboration · 2 authors totalFirst Sagittarius A* Event Horizon Telescope Results. II. EHT and Multiwavelength Observations, Data Processing, and Calibration
The Astrophysical Journal Letters · DOI 10.3847/2041-8213/ac6675 · Source: iop+orcid+smithsonianGreg Lindahl, Event Horizon Telescope Collaboration · 2 authors totalFirst Sagittarius A* Event Horizon Telescope Results. I. The Shadow of the Supermassive Black Hole in the Center of the Milky Way
The Astrophysical Journal Letters · DOI 10.3847/2041-8213/ac6674 · Source: iop+orcid+smithsonianGreg Lindahl, Event Horizon Telescope Collaboration · 2 authors totalFirst Sagittarius A* Event Horizon Telescope Results. V. Testing Astrophysical Models of the Galactic Center Black Hole
The Astrophysical Journal Letters · DOI 10.3847/2041-8213/ac6672 · Source: iop+orcid+smithsonianGreg Lindahl, Event Horizon Telescope Collaboration · 2 authors totalFirst Sagittarius A* Event Horizon Telescope Results. III. Imaging of the Galactic Center Supermassive Black Hole
The Astrophysical Journal Letters · DOI 10.3847/2041-8213/ac6429 · Source: iop+orcid+smithsonianGreg Lindahl, Event Horizon Telescope Collaboration · 2 authors totalResolving the Inner Parsec of the Blazar J1924-2914 with the Event Horizon Telescope
The Astrophysical Journal · DOI 10.3847/1538-4357/ac7a40 · Source: iop+orcid+smithsonianGreg Lindahl, Event Horizon Telescope Collaboration · 2 authors totalScalable Artificial Intelligence for Earth Observation Data Using Hopsworks
Remote. Sens. · DOI 10.3390/RS14081889 · Source: dblp+first-party-career-authorityJim Dowling, Desta Haileselassie Hagos, Theofilos Kakantousis, Sina Sheikholeslami, Tianze Wang, Vladimir Vlassov, Amir Hossein Payberah, Moritz Meister · 9 authors totalPediatric Sarcomas: The Next Generation of Molecular Studies
Cancers · DOI 10.3390/cancers14102515 · 2 citations · Source: openalex+authoritative-profilePetros Giannikopoulos, David M. Parham · 2 authors totalCharacterizing Dysarthria Diversity for Automatic Speech Recognition: A Tutorial From the Clinical Perspective.
Frontiers Comput. Sci. · DOI 10.3389/fcomp.2022.770210 · Source: dblpKatrin Tomanek, Hannah P. Rowe, Sarah E. Gutz, Marc F. Maffei, Jordan R. Green · 5 authors totalA Universal Screening Tool for Dyslexia by a Web-Game and Machine Learning
Frontiers in Computer Science · DOI 10.3389/fcomp.2021.628634 · 27 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Maria Rauschenberger, Ricardo Baeza‐Yates, Luz Rello · 4 authors totalSomething new and different: The Unified Medical Language System
Information Services & Use · DOI 10.3233/isu-210138 · 9 citations · Source: openalex+first-party-career-authorityMark Samuel Tuttle, Betsy L. Humphreys, Mark S. Tuttle · 3 authors totalQuantum Goemans-Williamson Algorithm with the Hadamard Test and Approximate Amplitude Constraints
Quantum · DOI 10.22331/q-2023-07-12-1057 · arXiv 2206.14999 · 13 citations · Source: semantic-scholarSemidefinite programs are optimization methods with a wide array of applications, such as approximating difficult combinatorial problems. One such semidefinite program is the Goemans-Williamson algorithm, a popular integer relaxation technique. We introduce a variational quantum algorithm for the Goemans-Williamson algorithm that uses only n+1 qubits, a constant number of circuit preparations, and poly(n) expectation values in order to approximately solve semidefinite programs with up to N=2n variables and M∼O(N) constraints. Efficient optimization is achieved by encoding the objective matrix as a properly parameterized unitary conditioned on an auxilary qubit, a technique known as the Hadamard Test. The Hadamard Test enables us to optimize the objective function by estimating only a single expectation value of the ancilla qubit, rather than separately estimating exponentially many expectation values. Similarly, we illustrate that the semidefinite programming constraints can be effectively enforced by implementing a second Hadamard Test, as well as imposing a polynomial number of Pauli string amplitude constraints. We demonstrate the effectiveness of our protocol by devising an efficient quantum implementation of the Goemans-Williamson algorithm for various NP-hard problems, including MaxCut. Our method exceeds the performance of analogous classical methods on a diverse subset of well-studied MaxCut problems from the GSet library.
Jean Kossaifi, T. Patti, Anima Anandkumar, S. Yelin · 4 authors totalCharacterizing the Discourse of Popular Diets to Describe Information Dispersal and Identify Leading Voices, Interaction, and Themes of Mental Health: Social Network Analysis (Preprint)
DOI 10.2196/preprints.38245 · 0 citations · Source: openalex+first-party-career-authorityMarc Smith, Melissa Eaton, Yasmine Probst, Marc A. Smith · 4 authors totalProsodic characteristics of prepausal words produced by patients with neurodegenerative disease
Speech Prosody 2022 · DOI 10.21437/speechprosody.2022-25 · 4 citations · Source: openalex+first-party-career-authorityMark Liberman, Sunghye Cho, Galit Agmon, Sanjana Shellikeri, Katheryn A Q Cousins, Sharon Ash, David J. Irwin, Meredith Spindler · 13 authors totalThe mapping between syntactic and prosodic phrasing in English and Mandarin
Interspeech 2022 · DOI 10.21437/interspeech.2022-10726 · 2 citations · Source: openalex+first-party-career-authorityMark Liberman, Jianjing Kuang, May Pik Yu Chan, Nari Rhee, Hongwei Ding · 5 authors totalContext-Aware Abbreviation Expansion Using Large Language Models.
NAACL-HLT · DOI 10.18653/v1/2022.naacl-main.91 · Source: dblpKatrin Tomanek, Shanqing Cai, Subhashini Venugopalan, Ajit Narayanan, Meredith Ringel Morris, Michael P. Brenner · 6 authors totalHarmless Transfer Learning for Item Embeddings
Findings of the Association for Computational Linguistics: NAACL 2022 · DOI 10.18653/v1/2022.findings-naacl.38 · 2 citations · Source: openalexLearning embedding layers (for classes, words, items, etc.) is a key component of lots of applications, ranging from natural language processing, recommendation systems to electronic health records, etc. However, the frequency of real-world items follows a long-tail distribution in these applications, causing naive training methods perform poorly on the rare items. A line of previous works address this problem by transferring the knowledge from the frequent items to rare items by introducing an auxiliary transfer loss. However, when defined improperly, the transfer loss may introduce harmful biases and deteriorate the performance.
Dhruv Choudhary, Chengyue Gong, Xiaocong Du, Bhargav Bhushanam, Qiang Liu, Arun Kejariwal · 6 authors totalImpact of Pretraining Term Frequencies on Few-Shot Numerical Reasoning
Conference on Empirical Methods in Natural Language Processing · DOI 10.18653/v1/2022.findings-emnlp.59 · 192 citations · Source: semantic-scholar,
Sameer Singh, Yasaman Razeghi, IV RobertL.Logan, Matt Gardner · 4 authors totalThe Aligned Multimodal Movie Treebank: An audio, video, dependency-parse treebank
Conference on Empirical Methods in Natural Language Processing · DOI 10.18653/v1/2022.emnlp-main.648 · 5 citations · Source: semantic-scholarTreebanks have traditionally included only text and were derived from written sources such as newspapers or the web. We introduce the Aligned Multimodal Movie Treebank (AMMT), an English language treebank derived from dialog in Hollywood movies which includes transcriptions of the audio-visual streams with word-level alignment, as well as part of speech tags and dependency parses in the Universal Dependencies formalism. AMMT consists of 31,264 sentences and 218,090 words, that will amount to the 3rd largest UD English treebank and the only multimodal treebank in UD. To help with the web-based annotation effort, we also introduce the Efficient Audio Alignment Annotator (EAAA), a companion tool that enables annotators to significantly speed-up their annotation processes.
Ignacio Cases, A. Yaari, Jan DeWitt, Henry Hu, Bennett Stankovits, Sue Felshin, Yevgeni Berzak, Helena Aparicio · 10 authors totalIdentifying stable speech-language markers of autism in children: Preliminary evidence from a longitudinal telephony-based study
DOI 10.18653/v1/2022.clpsych-1.4 · 3 citations · Source: openalex+first-party-career-authorityMark Liberman, Sunghye Cho, Riccardo Fusaroli, Maggie Rose Pelella, Kimberly Tena, Azia Knox, Aili Hauptmann, Maxine Covello · 19 authors totalDeep Learning for Recommender Systems: A Netflix Case Study
The AI Magazine · DOI 10.1609/aimag.v42i3.18140 · 168 citations · Source: semantic-scholar+openalexDeep learning has profoundly impacted many areas of machine learning. However, it took a while for its impact to be felt in the field of recommender systems. In this article, we outline some of the challenges encountered and lessons learned in using deep learning for recommender systems at Netflix. We first provide an overview of the various recommendation tasks on the Netflix service. We found that different model architectures excel at different tasks. Even though many deep-learning models can be understood as extensions of existing (simple) recommendation algorithms, we initially did not observe significant improvements in performance over well-tuned non-deep-learning approaches. Only when we added numerous features of heterogeneous types to the input data, deep-learning models did start to shine in our setting. We also observed that deep-learning methods can exacerbate the problem of offline–online metric (mis-)alignment. After addressing these challenges, deep learning has ultimately resulted in large improvements to our recommendations as measured by both offline and online metrics. On the practical side, integrating deep-learning toolboxes in our system has made it faster and easier to implement and experiment with both deep-learning and non-deep-learning approaches for various recommendation tasks. We conclude this article by summarizing our take-aways that may generalize to other applications beyond Netflix.
Justin Basilico, Harald Steck, Linas Baltrunas, Ehtsham Elahi, Dawen Liang, Yves Raimond · 6 authors totalSimilarity Search for Efficient Active Learning and Search of Rare Concepts
Proceedings of the AAAI Conference on Artificial Intelligence · DOI 10.1609/aaai.v36i6.20591 · 22 citations · Source: openalex+authoritative-profilePeter Bailis, Cody Coleman, Edward Chou, Julian Katz-Samuels, Sean Chang Culatana, Alexander C. Berg, Robert D. Nowak, Roshan Sumbaly · 10 authors totalKeep CALM and CRDT On
Proceedings of the VLDB Endowment · DOI 10.14778/3574245.3574268 · arXiv 2210.12605 · 14 citations · Source: semantic-scholarDespite decades of research and practical experience, developers have few tools for programming reliable distributed applications without resorting to expensive coordination techniques. Conflict-free replicated datatypes (CRDTs) are a promising line of work that enable coordination-free replication and offer certain eventual consistency guarantees in a relatively simple object-oriented API. Yet CRDT guarantees extend only to data updates; observations of CRDT state are unconstrained and unsafe. We propose an agenda that embraces the simplicity of CRDTs, but provides richer, more uniform guarantees. We extend CRDTs with a query model that reasons about which queries are safe without coordination by applying monotonicity results from the CALM Theorem, and lay out a larger agenda for developing CRDT data stores that let developers safely and efficiently interact with replicated application state.
Shadaj Laddad, Conor Power, Mae Milano, Alvin Cheung, Natacha Crooks, Joseph M. Hellerstein · 6 authors totalParallelism-Optimizing Data Placement for Faster Data-Parallel Computations
Proceedings of the VLDB Endowment · DOI 10.14778/3574245.3574260 · 5 citations · Source: openalex+authoritative-profilePeter Bailis, Nirvik Baruah, Peter Kraft, Fiodar Kazhamiaka, Matei Zaharia · 5 authors totalManu: A Cloud Native Vector Database Management System
PVLDB 15(12) (VLDB 2022) · DOI 10.14778/3554821.3554843 · arXiv 2206.13843 · Source: arxiv+crossrefWith the development of learning-based embedding models, embedding vectors are widely used for analyzing and searching unstructured data. As vector collections exceed billion-scale, fully managed and horizontally scalable vector databases are necessary. In the past three years, through interaction with our 1200+ industry users, we have sketched a vision for the features that next-generation vector databases should have, which include long-term evolvability, tunable consistency, good elasticity, and high performance. We present Manu, a cloud native vector database that implements these features. It is difficult to integrate all these features if we follow traditional DBMS design rules. As most vector data applications do not require complex data models and strong data consistency, our design philosophy is to relax the data model and consistency constraints in exchange for the aforementioned features. Specifically, Manu firstly exposes the write-ahead log (WAL) and binlog as backbone services. Secondly, write components are designed as log publishers while all read-only analytic and search components are designed as independent subscribers to the log services. Finally, we utilize multi-version concurrency control (MVCC) and a delta consistency model to simplify the communication and cooperation among the system components. These designs achieve a low coupling among the system components, which is essential for elasticity and evolution. We also extensively optimize Manu for performance and usability with hardware-aware implementations and support for complex search semantics.
James Luan, Rentong Guo, Xiaofan Luan, Long Xiang, Xiao Yan, Xiaomeng Yi, Jigao Luo, Qianya Cheng · 15 authors totalTAOBench
Proceedings of the VLDB Endowment · DOI 10.14778/3538598.3538616 · 17 citations · Source: openalex+authoritative-profilePeter Bailis, Audrey Cheng, Xiao Shi, Aaron Kabcenell, Shilpa Lawande, Hamza Qadeer, Jason Chan, Harrison Tin · 13 authors totalLexical and Acoustic Speech Features Relating to Alzheimer Disease Pathology
Neurology · DOI 10.1212/wnl.0000000000200581 · 50 citations · Source: openalex+first-party-career-authorityMark Liberman, Sunghye Cho, Katheryn A Q Cousins, Sanjana Shellikeri, Sharon Ash, David J. Irwin, Murray Grossman, Naomi Nevler · 8 authors totalIntroduction
Annual Review of Linguistics · DOI 10.1146/annurev-li-08-120721-100001 · 0 citations · Source: openalex+first-party-career-authorityMark Liberman, Colin Phillips · 2 authors totalReport on the 12th Temporal Web Analytics Workshop (TempWeb 2022) at WWW 2022
ACM SIGIR Forum · DOI 10.1145/3582900.3582909 · 0 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Marc Spaniol, Ricardo Baeza‐Yates, Omar Alonso · 4 authors totalTwHIN-BERT: A Socially-Enriched Pre-trained Language Model for Multilingual Tweet Representations at Twitter
Knowledge Discovery and Data Mining · DOI 10.1145/3580305.3599921 · arXiv 2209.07562 · 132 citations · Source: semantic-scholarPre-trained language models (PLMs) are fundamental for natural language processing applications. Most existing PLMs are not tailored to the noisy user-generated text on social media, and the pre-training does not factor in the valuable social engagement logs available in a social network. We present TwHIN-BERT, a multilingual language model productionized at Twitter, trained on in-domain data from the popular social network. TwHIN-BERT differs from prior pre-trained language models as it is trained with not only text-based self-supervision but also with a social objective based on the rich social engagements within a Twitter heterogeneous information network (TwHIN). Our model is trained on 7 billion tweets covering over 100 distinct languages, providing a valuable representation to model short, noisy, user-generated text. We evaluate our model on various multilingual social recommendation and semantic understanding tasks and demonstrate significant metric improvement over established pre-trained language models. We open-source TwHIN-BERT and our curated hashtag prediction and social engagement benchmark datasets to the research community.
Yury Malkov, Xinyang Zhang, Omar U. Florez, Serim Park, B. McWilliams, Jiawei Han, Ahmed El-Kishky · 7 authors totalFHIR: Reducing Friction in the Exchange of Healthcare Data
ACM Queue · DOI 10.1145/3565861 · 0 citations · Source: semantic-scholarPat Helland, Jim Agnew, Adam Cole · 3 authors totalAccelerating automation of digital health applications via cloud native approach
WOC@Middleware · DOI 10.1145/3565384.3565888 · 2 citations · Source: semantic-scholarWriting thread safe code for concurrent processing requires experience and training, thus legacy research code are usually single threaded, which post a challenge when it comes to scaling. This challenge is harder to overcome in the digital health domain, where code changes might trigger a regulatory review process. In this work, we report a solution of leveraging container technology to convert single threaded legacy code into cloud native services, scaling out data processing throughput via data parallelism. We tested the setup with a batch data processing job on the EverythingALS dataset and obtained 8X speed up compared to single threaded processing. This solution is built as part of IBM Health Guardian, a digital health tool suite. It is generalizable and can be adapted to other projects. The work greatly improves the automation of adoption of legacy research code in the evolving digital health domain. It will attract more open domain research contributions.
Dean Wampler, Bo Wen, Y. Koyfman, Hongfei Tian, B. Lublinsky, R. Norel, Carla Agurto, D. Wampler · 8 authors totalUnstoppable DAOs for web3 disruption
DOI 10.1145/3565383.3566112 · 4 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Rowdy Chotkan, Jérémie Decouchant · 3 authors totalKatara: synthesizing CRDTs with verified lifting
OOPSLA (PACMPL) · DOI 10.1145/3563336 · arXiv 2205.12425 · 27 citations · Source: semantic-scholarConflict-free replicated data types (CRDTs) are a promising tool for designing scalable, coordination-free distributed systems. However, constructing correct CRDTs is difficult, posing a challenge for even seasoned developers. As a result, CRDT development is still largely the domain of academics, with new designs often awaiting peer review and a manual proof of correctness. In this paper, we present Katara, a program synthesis-based system that takes sequential data type implementations and automatically synthesizes verified CRDT designs from them. Key to this process is a new formal definition of CRDT correctness that combines a reference sequential type with a lightweight ordering constraint that resolves conflicts between non-commutative operations. Our process follows the tradition of work in verified lifting, including an encoding of correctness into SMT logic using synthesized inductive invariants and hand-crafted grammars for the CRDT state and runtime. Katara is able to automatically synthesize CRDTs for a wide variety of scenarios, from reproducing classic CRDTs to synthesizing novel designs based on specifications in existing literature. Crucially, our synthesized CRDTs are fully, automatically verified, eliminating entire classes of common errors and reducing the process of producing a new CRDT from a painstaking paper proof of correctness to a lightweight specification.
Shadaj Laddad, Conor Power, Mae Milano, Alvin Cheung, Joseph M. Hellerstein · 5 authors totalLearning UML database design and modeling with AutoER.
MoDELS (Companion) · DOI 10.1145/3550356.3559091 · Source: dblp+ubc-authorityRamon Lawrence, Sarah Foss, Tatiana Urazova · 3 authors totalRedShift: Transparent SNARKs from List Polynomial Commitments
ACM CCS · DOI 10.1145/3548606.3560657 · Source: acm-ccs+dblp+matter-labs-authorityAlexander Vlasov, Assimakis A. Kattis, Konstantin Panarin · 3 authors totalI'm Probably Less Deterministic Than I Used to Be
ACM Queue · DOI 10.1145/3546935 · 0 citations · Source: semantic-scholarPat Helland · 1 author totalMetadata-based retrieval for resolution recommendation in AIOps
ESEC/SIGSOFT FSE · DOI 10.1145/3540250.3558964 · 3 citations · Source: crossref+semantic-scholarRuchi Mahindru, Harshit Kumar, R. Mahindru, Debanjana Kar · 4 authors totalTwHIN: Embedding the Twitter Heterogeneous Information Network for Personalized Recommendation
Knowledge Discovery and Data Mining · DOI 10.1145/3534678.3539080 · arXiv 2202.05387 · 75 citations · Source: semantic-scholarSocial networks, such as Twitter, form a heterogeneous information network (HIN) where nodes represent domain entities (e.g., user, content, advertiser, etc.) and edges represent one of many entity interactions (e.g, a user re-sharing content or "following" another). Interactions from multiple relation types can encode valuable information about social network entities not fully captured by a single relation; for instance, a user's preference for accounts to follow may depend on both user-content engagement interactions and the other users they follow. In this work, we investigate knowledge-graph embeddings for entities in the Twitter HIN (TwHIN); we show that these pretrained representations yield significant offline and online improvement for a diverse range of downstream recommendation and classification tasks: personalized ads rankings, account follow-recommendation, offensive content detection, and search ranking. We discuss design choices and practical challenges of deploying industry-scale HIN embeddings, including compressing them to reduce end-to-end model latency and handling parameter drift across versions.
Yury Malkov, Ahmed El-Kishky, Thomas Markovich, Serim Park, C. Verma, Baekjin Kim, R. Eskander, Frank Portman · 11 authors totalLooper: An End-to-End ML Platform for Product Decisions
Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining · DOI 10.1145/3534678.3539059 · 10 citations · Source: openalex+career-authoritySal Uryasev, Igor L. Markov, Hanson Wang, Nitya Kasturi, Shaun Singh, Mia R. Garrard, Yin Huang, Sze Wai Yuen · 19 authors totalAutoShard: Automated Embedding Table Sharding for Recommender Systems
Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining · DOI 10.1145/3534678.3539034 · 23 citations · Source: openalexEmbedding learning is an important technique in deep recommendation models to map categorical features to dense vectors. However, the embedding tables often demand an extremely large number of parameters, which become the storage and efficiency bottlenecks. Distributed training solutions have been adopted to partition the embedding tables into multiple devices. However, the embedding tables can easily lead to imbalances if not carefully partitioned. This is a significant design challenge of distributed systems named embedding table sharding, i.e., how we should partition the embedding tables to balance the costs across devices, which is a non-trivial task because 1) it is hard to efficiently and precisely measure the cost, and 2) the partition problem is known to be NP-hard. In this work, we introduce our novel practice in Meta, namely AutoShard, which uses a neural cost model to directly predict the multi-table costs and leverages deep reinforcement learning to solve the partition problem. Experimental results on an open-sourced large-scale synthetic dataset and Meta's production dataset demonstrate the superiority of AutoShard over the heuristics. Moreover, the learned policy of AutoShard can transfer to sharding tasks with various numbers of tables and different ratios of the unseen tables without any fine-tuning. Furthermore, AutoShard can efficiently shard hundreds of tables in seconds. The effectiveness, transferability, and efficiency of AutoShard make it desirable for production use. Our algorithms have been deployed in Meta production environment. A prototype is available at https://github.com/daochenzha/autoshard
Dhruv Choudhary, Daochen Zha, Louis Feng, Bhargav Bhushanam, Jade Nie, Yuandong Tian, Jay Chae, Yinbin Ma · 10 authors totalMapping of Financial Services datasets using Human-in-the-Loop
International Conference on AI in Finance · DOI 10.1145/3533271.3561705 · 2 citations · Source: crossref+semantic-scholarIncreasing access to financial services data helps accelerate the monitoring and management of datasets and facilitates better business decision-making. However, financial services datasets are typically vast, ranging in terabytes of data, containing both structured and unstructured. It is a laborious task to comb through all the data and map them reasonably. Mapping the data is important to perform comprehensive analysis and take informed business decisions. Based on client engagements, we have observed that there is a lack of industry standards for definitions of key terms and a lack of governance for maintaining business processes. This typically leads to disconnected siloed datasets generated from disintegrated systems. To address these challenges, we developed a novel methodology DaME (Data Mapping Engine) that performs data mapping by training a data mapping engine and utilizing human-in-the-loop techniques. The results from the industrial application and evaluation of DaME on a financial services dataset are encouraging that it can help reduce manual effort by automating data mapping and reusing the learning. The accuracy from our dataset in the application is much higher at 69% compared to the existing state-of-the-art with an accuracy of 34%. It has also helped improve the productivity of the industry practitioners, by saving them 14,000 hours of time spent manually mapping vast data stores over a period of ten months.
Ruchi Mahindru, Shubhi Asthana, R. Mahindru · 3 authors totalPredictability and Surprise in Large Generative Models
Conference on Fairness, Accountability and Transparency · DOI 10.1145/3531146.3533229 · arXiv 2202.07785 · 374 citations · Source: semantic-scholarLarge-scale pre-training has recently emerged as a technique for creating capable, general-purpose, generative models such as GPT-3, Megatron-Turing NLG, Gopher, and many others. In this paper, we highlight a counterintuitive property of such models and discuss the policy implications of this property. Namely, these generative models have a paradoxical combination of predictable loss on a broad training distribution (as embodied in their ”scaling laws”), and unpredictable specific capabilities, inputs, and outputs. We believe that the high-level predictability and appearance of useful capabilities drives rapid development of such models, while the unpredictable qualities make it difficult to anticipate the consequences of model deployment. We go through examples of how this combination can lead to socially harmful behavior with examples from the literature and real world observations, and we also perform two novel experiments to illustrate our point about harms from unpredictability. Furthermore, we analyze how these conflicting properties combine to give model developers various motivations for deploying these models, and challenges that can hinder deployment. We conclude with a list of possible interventions the AI community may take to increase the chance of these models having a beneficial impact. We intend for this paper to be useful to policymakers who want to understand and regulate AI systems, technologists who care about the potential policy impact of their work, funders who want to support work addressing these challenges, and academics who want to analyze, critique, and potentially develop large generative models.
Tom Brown, Deep Ganguli, Danny Hernandez, Liane Lovitt, Nova Dassarma, T. Henighan, Andy Jones, Nicholas Joseph · 30 authors total