Papers.
Research connected to its authors, projects, companies, talks, events, and the rest of the graph.
Add a paper ↗Web3Recommend: Decentralised recommendations with trust and relevance
arXiv (Cornell University) · DOI 10.48550/arxiv.2307.01411 · 1 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Rohan Madhwal · 2 authors totalBeyond NTK with Vanilla Gradient Descent: A Mean-Field Analysis of Neural Networks with Polynomial Width, Samples, and Time.
CoRR · DOI 10.48550/arXiv.2306.16361 · arXiv 2306.16361 · Source: dblp+stanford-authorityTengyu Ma, Arvind V. Mahankali, Jeff Z. HaoChen, Kefan Dong, Margalit Glasgow, Tengyu Ma 0001 · 6 authors totalTowards Sybil Resilience in Decentralized Learning
arXiv (Cornell University) · DOI 10.48550/arxiv.2306.15044 · 0 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Thomas Werthenbach · 2 authors totalHuman-AI Coevolution
arXiv (Cornell University) · DOI 10.48550/arxiv.2306.13723 · 9 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Dino Pedreschi, Luca Pappalardo, Ferragina, Emanuele, Ricardo Baeza‐Yates, Albert-Ĺaszló Barabási, Frank Dignum, Virginia Dignum · 17 authors totalThe Inductive Bias of Flatness Regularization for Deep Matrix Factorization.
CoRR · DOI 10.48550/arXiv.2306.13239 · arXiv 2306.13239 · Source: dblp+stanford-authorityTengyu Ma, Khashayar Gatmiry, Zhiyuan Li 0005, Ching-Yao Chuang, Sashank J. Reddi, Tengyu Ma 0001, Stefanie Jegelka · 7 authors totalOptimizing Stateful Dataflow with Local Rewrites
arXiv preprint · arXiv 2306.10585 · 3 citations · Source: semantic-scholarOptimizing a stateful dataflow language is a challenging task. There are strict correctness constraints for preserving properties expected by downstream consumers, a large space of possible optimizations, and complex analyses that must reason about the behavior of the program over time. Classic compiler techniques with specialized optimization passes yield unpredictable performance and have complex correctness proofs. But with e-graphs, we can dramatically simplify the process of building a correct optimizer while yielding more consistent results! In this short paper, we discuss our early work using e-graphs to develop an optimizer for a the Hydroflow dataflow language. Our prototype demonstrates that composing simple, easy-to-prove rewrite rules is sufficient to match techniques in hand-optimized systems.
Shadaj Laddad, Conor Power, Tyler Hou, Alvin Cheung, Joseph M. Hellerstein · 5 authors totalA Heavy-Tailed Algebra for Probabilistic Programming
NeurIPS 2023 · DOI 10.48550/arXiv.2306.09262 · arXiv 2306.09262 · 4 citations · Source: semanticscholar+arxivThe authors propose a systematic static-analysis approach that reasons about the tails of random variables flowing through a probabilistic program, developing an algebra for the tail behaviour of distributions under common operations and using it to inform proposal and variational family design in probabilistic programming systems.
Feynman Liang, Liam Hodgkinson, Michael W. Mahoney · 3 authors totalAnticipatory Music Transformer
Trans. Mach. Learn. Res. · DOI 10.48550/arXiv.2306.08620 · arXiv 2306.08620 · 46 citations · Source: semantic-scholarWe introduce anticipation: a method for constructing a controllable generative model of a temporal point process (the event process) conditioned asynchronously on realizations of a second, correlated process (the control process). We achieve this by interleaving sequences of events and controls, such that controls appear following stopping times in the event sequence. This work is motivated by problems arising in the control of symbolic music generation. We focus on infilling control tasks, whereby the controls are a subset of the events themselves, and conditional generation completes a sequence of events given the fixed control events. We train anticipatory infilling models using the large and diverse Lakh MIDI music dataset. These models match the performance of autoregressive models for prompted music generation, with the additional capability to perform infilling control tasks, including accompaniment. Human evaluators report that an anticipatory model produces accompaniments with similar musicality to even music composed by humans over a 20-second clip.
David Hall, John Thickstun, David Leo Wright Hall, Chris Donahue, Percy Liang · 5 authors totalh2oGPT: Democratizing Large Language Models
arXiv · DOI 10.48550/arXiv.2306.08161 · arXiv 2306.08161 · 9 citations · Source: openalex+arxivApplications built on top of Large Language Models (LLMs) such as GPT-4 represent a revolution in AI due to their human-level capabilities in natural language processing. However, they also pose many significant risks such as the presence of biased, private, or harmful text, and the unauthorized inclusion of copyrighted material. We introduce h2oGPT, a suite of open-source code repositories for the creation and use of LLMs based on Generative Pretrained Transformers (GPTs). The goal of this project is to create the world's best truly open-source alternative to closed-source approaches. In collaboration with and as part of the incredible and unstoppable open-source community, we open-source several fine-tuned h2oGPT models from 7 to 40 Billion parameters, ready for commercial use under fully permissive Apache 2.0 licenses. Included in our release is 100\% private document search using natural language. Open-source language models help boost AI development and make it more accessible and trustworthy. They lower entry hurdles, allowing people and groups to tailor these models to their needs. This openness increases innovation, transparency, and fairness. An open-source strategy is needed to share AI benefits fairly, and H2O.ai will continue to democratize AI and LLMs.
Arno Candel, SriSatish Ambati, Jon McKinney, Philipp Singer, Pascal Pfeiffer, Maximilian Jeblick, Prithvi Prabhu, Jeff Gambera · 15 authors totalFormalizing Box Inference for Capture Calculus
arXiv 2306.06496 · 6 citations · Source: arxivCapture calculus has recently been proposed as a solution to effect checking, achieved by tracking the captured references of terms in the types. Boxes, along with the box and unbox operations, are a crucial construct in capture calculus, which maintains the hygiene of types and improves the expressiveness of polymorphism over capturing types. Despite their usefulness in the formalism, boxes would soon become a heavy notational burden for users when the program grows. It thus necessitates the inference of boxes when integrating capture checking into a mainstream programming language. In this paper, we develop a formalisation of box inference for capture calculus. We begin by introducing a semi-algorithmic variant of the capture calculus, from which we derive an inference system where typed transformations are applied to complete missing box operations in programs. Then, we propose a type-level system that performs provably equivalent inference on the type level, without actually transforming the program. In the metatheory, we establish the relationships between these systems and capture calculus, thereby proving both the soundness and the completeness of box inference.
Martin Odersky, Yichen Xu · 2 authors totalEvaluating the Social Impact of Generative AI Systems in Systems and Society
arXiv · arXiv 2306.05949 · Source: arxiv+dblp+hugging-face-career-authorityIrene Solaiman, Zeerak Talat, William Agnew, and collaborators · 4 authors totalFair multilingual vandalism detection system for Wikipedia
arXiv (Cornell University) · DOI 10.48550/arxiv.2306.01650 · 1 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Mykola Trokhymovych, Muniza Aslam, Ai-Jou Chou, Ricardo Baeza‐Yates, Diego Sáez-Trumper · 6 authors totalKarma: Resource Allocation for Dynamic Demands
USENIX Symposium on Operating Systems Design and Implementation · arXiv 2305.17222 · 27 citations · Source: semantic-scholarWe consider the problem of fair resource allocation in a system where user demands are dynamic, that is, where user demands vary over time. Our key observation is that the classical max-min fairness algorithm for resource allocation provides many desirable properties (e.g., Pareto efficiency, strategy-proofness, and fairness), but only under the strong assumption of user demands being static over time. For the realistic case of dynamic user demands, the max-min fairness algorithm loses one or more of these properties. We present Karma, a new resource allocation mechanism for dynamic user demands. The key technical contribution in Karma is a credit-based resource allocation algorithm: in each quantum, users donate their unused resources and are assigned credits when other users borrow these resources; Karma carefully orchestrates the exchange of credits across users (based on their instantaneous demands, donated resources and borrowed resources), and performs prioritized resource allocation based on users' credits. We theoretically establish Karma guarantees related to Pareto efficiency, strategy-proofness, and fairness for dynamic user demands. Empirical evaluations over production workloads show that these properties translate well into practice: Karma is able to reduce disparity in performance across users to a bare minimum while maintaining Pareto-optimal system-wide performance.
Anurag Khandelwal, Midhul Vuppalapati, Giannis Fikioris, Rachit Agarwal, Asaf Cidon, Éva Tardos · 6 authors totalMeta-MeTTa: an operational semantics for MeTTa
CoRR · DOI 10.48550/arXiv.2305.17218 · Source: dblp+arxiv+career-authorityLucius Gregory Meredith, Ben Goertzel, Jonathan Warrell, Adam Vandervorst · 4 authors totalLarge Language Models as Tool Makers.
CoRR · DOI 10.48550/arXiv.2305.17126 · arXiv 2305.17126 · Source: dblp+stanford-authorityTengyu Ma, Tianle Cai, Xuezhi Wang 0002, Tengyu Ma 0001, Xinyun Chen, Denny Zhou · 6 authors totalRevisiting non-English Text Simplification: A Unified Multilingual Benchmark
ACL · DOI 10.48550/arXiv.2305.15678 · arXiv 2305.15678 · 45 citations · Source: arxiv+semantic-scholarRecent advancements in high-quality, large-scale English resources have pushed the frontier of English Automatic Text Simplification (ATS) research. However, less work has been done on multilingual text simplification due to the lack of a diverse evaluation benchmark that covers complex-simple sentence pairs in many languages. This paper introduces the MultiSim benchmark, a collection of 27 resources in 12 distinct languages containing over 1.7 million complex-simple sentence pairs. This benchmark will encourage research in developing more effective multilingual text simplification models and evaluation metrics. Our experiments using MultiSim with pre-trained multilingual language models reveal exciting performance improvements from multilingual training in non-English settings. We observe strong performance from Russian in zero-shot cross-lingual transfer to low-resource languages. We further show that few-shot prompting with BLOOM-176b achieves comparable quality to reference simplifications outperforming fine-tuned models in most languages. We validate these findings through human evaluation.
Michael Ryan, Michael J. Ryan, Tarek Naous, Wei Xu · 4 authors totalHaving Beer after Prayer? Measuring Cultural Bias in Large Language Models
ACL · DOI 10.48550/arXiv.2305.14456 · arXiv 2305.14456 · 211 citations · Source: arxiv+semantic-scholarAs the reach of large language models (LMs) expands globally, their ability to cater to diverse cultural contexts becomes crucial. Despite advancements in multilingual capabilities, models are not designed with appropriate cultural nuances. In this paper, we show that multilingual and Arabic monolingual LMs exhibit bias towards entities associated with Western culture. We introduce CAMeL, a novel resource of 628 naturally-occurring prompts and 20,368 entities spanning eight types that contrast Arab and Western cultures. CAMeL provides a foundation for measuring cultural biases in LMs through both extrinsic and intrinsic evaluations. Using CAMeL, we examine the cross-cultural performance in Arabic of 16 different LMs on tasks such as story generation, NER, and sentiment analysis, where we find concerning cases of stereotyping and cultural unfairness. We further test their text-infilling performance, revealing the incapability of appropriate adaptation to Arab cultural contexts. Finally, we analyze 6 Arabic pre-training corpora and find that commonly used sources such as Wikipedia may not be best suited to build culturally aware LMs, if used as they are without adjustment. We will make CAMeL publicly available at: https://github.com/tareknaous/camel
Michael Ryan, Tarek Naous, Michael J. Ryan, Alan Ritter, Wei Xu · 5 authors totalAlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
Neural Information Processing Systems · DOI 10.48550/arXiv.2305.14387 · arXiv 2305.14387 · 927 citations · Source: semantic-scholarLarge language models (LLMs) such as ChatGPT have seen widespread adoption due to their strong instruction-following abilities. Developing these LLMs involves a complex yet poorly understood workflow requiring training with human feedback. Replicating and understanding this instruction-following requires tackling three major challenges: the high cost of data collection, the lack of trustworthy evaluation, and the absence of reference method implementations. We address these challenges with AlpacaFarm, a simulator that enables research and development for learning from feedback at a low cost. First, we design LLM prompts to simulate human feedback that are 50x cheaper than crowdworkers and display high agreement with humans. Second, we propose an automatic evaluation and validate it against human instructions obtained on real-world interactions. Third, we contribute reference implementations for several methods (PPO, DPO, best-of-n, expert iteration, and more) that learn from pairwise feedback. Finally, as an end-to-end validation of AlpacaFarm, we train and evaluate eleven models on 10k pairs of real human feedback and show that rankings of models trained in AlpacaFarm match rankings of models trained on human data. As a demonstration of the research possible in AlpacaFarm, we find that methods that use a reward model can substantially improve over supervised fine-tuning and that our reference PPO implementation leads to a +10% improvement in win-rate against Davinci003. We release all components of AlpacaFarm at https://github.com/tatsu-lab/alpaca_farm.
Carlos Guestrin, Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Percy Liang · 9 authors totalSophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
International Conference on Learning Representations · DOI 10.48550/arXiv.2305.14342 · arXiv 2305.14342 · 307 citations · Source: semantic-scholarGiven the massive cost of language model pre-training, a non-trivial improvement of the optimization algorithm would lead to a material reduction on the time and cost of training. Adam and its variants have been state-of-the-art for years, and more sophisticated second-order (Hessian-based) optimizers often incur too much per-step overhead. In this paper, we propose Sophia, Second-order Clipped Stochastic Optimization, a simple scalable second-order optimizer that uses a light-weight estimate of the diagonal Hessian as the pre-conditioner. The update is the moving average of the gradients divided by the moving average of the estimated Hessian, followed by element-wise clipping. The clipping controls the worst-case update size and tames the negative impact of non-convexity and rapid change of Hessian along the trajectory. Sophia only estimates the diagonal Hessian every handful of iterations, which has negligible average per-step time and memory overhead. On language modeling with GPT models of sizes ranging from 125M to 1.5B, Sophia achieves a 2x speed-up compared to Adam in the number of steps, total compute, and wall-clock time, achieving the same perplexity with 50% fewer steps, less total compute, and reduced wall-clock time. Theoretically, we show that Sophia, in a much simplified setting, adapts to the heterogeneous curvatures in different parameter dimensions, and thus has a run-time bound that does not depend on the condition number of the loss.
David Hall, Tengyu Ma, Hong Liu, Zhiyuan Li, David Leo Wright Hall, Percy Liang · 6 authors totalUniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild
Neural Information Processing Systems · DOI 10.48550/arXiv.2305.11147 · arXiv 2305.11147 · 244 citations · Source: arxiv+semantic-scholarAchieving machine autonomy and human control often represent divergent objectives in the design of interactive AI systems. Visual generative foundation models such as Stable Diffusion show promise in navigating these goals, especially when prompted with arbitrary languages. However, they often fall short in generating images with spatial, structural, or geometric controls. The integration of such controls, which can accommodate various visual conditions in a single unified model, remains an unaddressed challenge. In response, we introduce UniControl, a new generative foundation model that consolidates a wide array of controllable condition-to-image (C2I) tasks within a singular framework, while still allowing for arbitrary language prompts. UniControl enables pixel-level-precise image generation, where visual conditions primarily influence the generated structures and language prompts guide the style and context. To equip UniControl with the capacity to handle diverse visual conditions, we augment pretrained text-to-image diffusion models and introduce a task-aware HyperNet to modulate the diffusion models, enabling the adaptation to different C2I tasks simultaneously. Trained on nine unique C2I tasks, UniControl demonstrates impressive zero-shot generation abilities with unseen visual conditions. Experimental results show that UniControl often surpasses the performance of single-task-controlled methods of comparable model sizes. This control versatility positions UniControl as a significant advancement in the realm of controllable visual generation.
Ran Xu, Can Qin, Shu Zhang, Ning Yu, Yihao Feng, Xinyi Yang, Yingbo Zhou, Haiquan Wang · 13 authors totalDoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining.
CoRR · DOI 10.48550/arXiv.2305.10429 · arXiv 2305.10429 · Source: dblp+stanford-authorityTengyu Ma, Sang Michael Xie, Hieu Pham 0001, Xuanyi Dong, Nan Du 0002, Hanxiao Liu, Yifeng Lu, Percy Liang · 10 authors totalPaLM 2 Technical Report
arXiv preprint · DOI 10.48550/arXiv.2305.10403 · arXiv 2305.10403 · 1,543 citations · Source: arxiv+semantic-scholarWe introduce PaLM 2, a new state-of-the-art language model that has better multilingual and reasoning capabilities and is more compute-efficient than its predecessor PaLM. PaLM 2 is a Transformer-based model trained using a mixture of objectives. Through extensive evaluations on English and multilingual language, and reasoning tasks, we demonstrate that PaLM 2 has significantly improved quality on downstream tasks across different model sizes, while simultaneously exhibiting faster and more efficient inference compared to PaLM. This improved efficiency enables broader deployment while also allowing the model to respond faster, for a more natural pace of interaction. PaLM 2 demonstrates robust reasoning capabilities exemplified by large improvements over PaLM on BIG-Bench and other reasoning tasks. PaLM 2 exhibits stable performance on a suite of responsible AI evaluations, and enables inference-time control over toxicity without additional overhead or impact on other capabilities. Overall, PaLM 2 achieves state-of-the-art performance across a diverse set of tasks and capabilities. When discussing the PaLM 2 family, it is important to distinguish between pre-trained models (of various sizes), fine-tuned variants of these models, and the user-facing products that use these models. In particular, user-facing products typically include additional pre- and post-processing steps. Additionally, the underlying models may evolve over time. Therefore, one should not expect the performance of user-facing products to exactly match the results reported in this report.
Brennan Saeta, Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri · 12 authors totalSymbol tuning improves in-context learning in language models.
CoRR · DOI 10.48550/arXiv.2305.08298 · arXiv 2305.08298 · Source: dblp+stanford-authorityTengyu Ma, Jerry W. Wei, Le Hou, Andrew K. Lampinen, Xiangning Chen, Da Huang, Yi Tay, Xinyun Chen · 11 authors totalFrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
Trans. Mach. Learn. Res. · DOI 10.48550/arXiv.2305.05176 · arXiv 2305.05176 · 870 citations · Source: semantic-scholarThere is a rapidly growing number of large language models (LLMs) that users can query for a fee. We review the cost associated with querying popular LLM APIs, e.g. GPT-4, ChatGPT, J1-Jumbo, and find that these models have heterogeneous pricing structures, with fees that can differ by two orders of magnitude. In particular, using LLMs on large collections of queries and text can be expensive. Motivated by this, we outline and discuss three types of strategies that users can exploit to reduce the inference cost associated with using LLMs: 1) prompt adaptation, 2) LLM approximation, and 3) LLM cascade. As an example, we propose FrugalGPT, a simple yet flexible instantiation of LLM cascade which learns which combinations of LLMs to use for different queries in order to reduce cost and improve accuracy. Our experiments show that FrugalGPT can match the performance of the best individual LLM (e.g. GPT-4) with up to 98% cost reduction or improve the accuracy over GPT-4 by 4% with the same cost. The ideas and findings presented here lay a foundation for using LLMs sustainably and efficiently.
Matei Zaharia, Lingjiao Chen, M. Zaharia, James Y. Zou · 4 authors totalLESS-VFL: Communication-Efficient Feature Selection for Vertical Federated Learning
International Conference on Machine Learning · arXiv 2305.02219 · 39 citations · Source: semantic-scholar+arxivWe propose LESS-VFL, a communication-efficient feature selection method for distributed systems with vertically partitioned data. We consider a system of a server and several parties with local datasets that share a sample ID space but have different feature sets. The parties wish to collaboratively train a model for a prediction task. As part of the training, the parties wish to remove unimportant features in the system to improve generalization, efficiency, and explainability. In LESS-VFL, after a short pre-training period, the server optimizes its part of the global model to determine the relevant outputs from party models. This information is shared with the parties to then allow local feature selection without communication. We analytically prove that LESS-VFL removes spurious features from model training. We provide extensive empirical evidence that LESS-VFL can achieve high accuracy and remove spurious features at a fraction of the communication cost of other feature selection approaches.
Nathalie Baracaldo, Timothy Castiglia, Yi Zhou, Shiqiang Wang, S. Kadhe, S. Patterson · 6 authors totalToward L-recovery of Nonlinear Functions: A Polynomial Sample Complexity Bound for Gaussian Random Fields.
CoRR · DOI 10.48550/arXiv.2305.00322 · arXiv 2305.00322 · Source: dblp+stanford-authorityTengyu Ma, Kefan Dong, Tengyu Ma 0001 · 3 authors totalDataComp: In search of the next generation of multimodal datasets
Neural Information Processing Systems · DOI 10.48550/arXiv.2304.14108 · arXiv 2304.14108 · 751 citations · Source: semantic-scholarMultimodal datasets are a critical component in recent breakthroughs such as Stable Diffusion and GPT-4, yet their design does not receive the same research attention as model architectures or training algorithms. To address this shortcoming in the ML ecosystem, we introduce DataComp, a testbed for dataset experiments centered around a new candidate pool of 12.8 billion image-text pairs from Common Crawl. Participants in our benchmark design new filtering techniques or curate new data sources and then evaluate their new dataset by running our standardized CLIP training code and testing the resulting model on 38 downstream test sets. Our benchmark consists of multiple compute scales spanning four orders of magnitude, which enables the study of scaling trends and makes the benchmark accessible to researchers with varying resources. Our baseline experiments show that the DataComp workflow leads to better training sets. In particular, our best baseline, DataComp-1B, enables training a CLIP ViT-L/14 from scratch to 79.2% zero-shot accuracy on ImageNet, outperforming OpenAI's CLIP ViT-L/14 by 3.7 percentage points while using the same training procedure and compute. We release DataComp and all accompanying code at www.datacomp.ai.
Vaishaal Shankar, S. Gadre, Gabriel Ilharco, Alex Fang, J. Hayase, G. Smyrnis, Thao Nguyen, Ryan Marten · 34 authors totalLatent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation
arXiv · DOI 10.48550/arxiv.2304.08477 · arXiv 2304.08477 · 151 citations · Source: semantic-scholarWe propose Latent-Shift -- an efficient text-to-video generation method based on a pretrained text-to-image generation model that consists of an autoencoder and a U-Net diffusion model. Learning a video diffusion model in the latent space is much more efficient than in the pixel space. The latter is often limited to first generating a low-resolution video followed by a sequence of frame interpolation and super-resolution models, which makes the entire pipeline very complex and computationally expensive. To extend a U-Net from image generation to video generation, prior work proposes to add additional modules like 1D temporal convolution and/or temporal attention layers. In contrast, we propose a parameter-free temporal shift module that can leverage the spatial U-Net as is for video generation. We achieve this by shifting two portions of the feature map channels forward and backward along the temporal dimension. The shifted features of the current frame thus receive the features from the previous and the subsequent frames, enabling motion learning without additional parameters. We show that Latent-Shift achieves comparable or better results while being significantly more efficient. Moreover, Latent-Shift can generate images despite being finetuned for T2V generation.
Sonal Gupta, Jie An, Songyang Zhang, Harry Yang, Jia-Bin Huang, Jiebo Luo, Xiaoyue Yin · 7 authors totalText-Conditional Contextualized Avatars For Zero-Shot Personalization
arXiv · DOI 10.48550/arxiv.2304.07410 · arXiv 2304.07410 · 4 citations · Source: semantic-scholarRecent large-scale text-to-image generation models have made significant improvements in the quality, realism, and diversity of the synthesized images and enable users to control the created content through language. However, the personalization aspect of these generative models is still challenging and under-explored. In this work, we propose a pipeline that enables personalization of image generation with avatars capturing a user's identity in a delightful way. Our pipeline is zero-shot, avatar texture and style agnostic, and does not require training on the avatar at all - it is scalable to millions of users who can generate a scene with their avatar. To render the avatar in a pose faithful to the given text prompt, we propose a novel text-to-3D pose diffusion model trained on a curated large-scale dataset of in-the-wild human poses improving the performance of the SOTA text-to-motion models significantly. We show, for the first time, how to leverage large-scale image datasets to learn human 3D pose parameters and overcome the limitations of motion capture datasets.
Sonal Gupta, S. Azadi, Thomas Hayes, Akbar Shah, Guan Pang, Devi Parikh · 6 authors totalUncovering Bias in Personal Informatics
arXiv (Cornell University) · DOI 10.48550/arxiv.2303.15592 · 0 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Sofia Yfantidou, Pavlos Sermpezis, Athena Vakali, Ricardo Baeza‐Yates · 5 authors totalAn Evaluation of Memory Optimization Methods for Training Neural Networks
arXiv 2303.14633 · 0 citations · Source: arxiv+semantic-scholarAs models continue to grow in size, the development of memory optimization methods (MOMs) has emerged as a solution to address the memory bottleneck encountered when training large models. To comprehensively examine the practical value of various MOMs, we have conducted a thorough analysis of existing literature from a systems perspective. Our analysis has revealed a notable challenge within the research community: the absence of standardized metrics for effectively evaluating the efficacy of MOMs. The scarcity of informative evaluation metrics hinders the ability of researchers and practitioners to compare and benchmark different approaches reliably. Consequently, drawing definitive conclusions and making informed decisions regarding the selection and application of MOMs becomes a challenging endeavor. To address the challenge, this paper summarizes the scenarios in which MOMs prove advantageous for model training. We propose the use of distinct evaluation metrics under different scenarios. By employing these metrics, we evaluate the prevailing MOMs and find that their benefits are not universal. We present insights derived from experiments and discuss the circumstances in which they can be advantageous.
Zhuohan Li, Xiaoxuan Liu, Siddharth Jha, Alvin Cheung · 4 authors totalART: Automatic multi-step reasoning and tool-use for large language models
arXiv.org · DOI 10.48550/arXiv.2303.09014 · arXiv 2303.09014 · 224 citations · Source: semantic-scholarLarge language models (LLMs) can perform complex reasoning in few- and zero-shot settings by generating intermediate chain of thought (CoT) reasoning steps. Further, each reasoning step can rely on external tools to support computation beyond the core LLM capabilities (e.g. search/running code). Prior work on CoT prompting and tool use typically requires hand-crafting task-specific demonstrations and carefully scripted interleaving of model generations with tool use. We introduce Automatic Reasoning and Tool-use (ART), a framework that uses frozen LLMs to automatically generate intermediate reasoning steps as a program. Given a new task to solve, ART selects demonstrations of multi-step reasoning and tool use from a task library. At test time, ART seamlessly pauses generation whenever external tools are called, and integrates their output before resuming generation. ART achieves a substantial improvement over few-shot prompting and automatic CoT on unseen tasks in the BigBench and MMLU benchmarks, and matches performance of hand-crafted CoT prompts on a majority of these tasks. ART is also extensible, and makes it easy for humans to improve performance by correcting errors in task-specific programs or incorporating new tools, which we demonstrate by drastically improving performance on select tasks with minimal human intervention.
Sameer Singh, Bhargavi Paranjape, Scott M. Lundberg, Hannaneh Hajishirzi, Luke Zettlemoyer, Marco Tulio Ribeiro · 6 authors totalHigh-throughput Generative Inference of Large Language Models with a Single GPU
ICML 2023 · DOI 10.48550/arXiv.2303.06865 · arXiv 2303.06865 · 856 citations · Source: arxiv+semantic-scholarThe high computational and memory requirements of large language model (LLM) inference make it feasible only with multiple high-end accelerators. Motivated by the emerging demand for latency-insensitive tasks with batched processing, this paper initiates the study of high-throughput LLM inference using limited resources, such as a single commodity GPU. We present FlexGen, a high-throughput generation engine for running LLMs with limited GPU memory. FlexGen can be flexibly configured under various hardware resource constraints by aggregating memory and computation from the GPU, CPU, and disk. By solving a linear programming problem, it searches for efficient patterns to store and access tensors. FlexGen further compresses the weights and the attention cache to 4 bits with negligible accuracy loss. These techniques enable FlexGen to have a larger space of batch size choices and thus significantly increase maximum throughput. As a result, when running OPT-175B on a single 16GB GPU, FlexGen achieves significantly higher throughput compared to state-of-the-art offloading systems, reaching a generation throughput of 1 token/s for the first time with an effective batch size of 144. On the HELM benchmark, FlexGen can benchmark a 30B model with a 16GB GPU on 7 representative sub-scenarios in 21 hours. The code is available at https://github.com/FMInference/FlexGen
Zhuohan Li, Ying Sheng, Lianmin Zheng, Binhang Yuan, Max Ryabinin, Daniel Y. Fu, Zhiqiang Xie, Beidi Chen · 14 authors totalLarger language models do in-context learning differently.
CoRR · DOI 10.48550/arXiv.2303.03846 · arXiv 2303.03846 · Source: dblp+stanford-authorityTengyu Ma, Jerry W. Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen · 11 authors totalDesigning a Fast and Flexible Quantum State Simulator
arXiv · arXiv 2303.01493 · Source: arxiv+career-authorityConstantin Gonciulea, Saveliy Yusufov, Charlee Stefanski · 3 authors totalBrainBERT: Self-supervised representation learning for intracranial recordings
International Conference on Learning Representations · DOI 10.48550/arXiv.2302.14367 · arXiv 2302.14367 · 102 citations · Source: semantic-scholarWe create a reusable Transformer, BrainBERT, for intracranial recordings bringing modern representation learning approaches to neuroscience. Much like in NLP and speech recognition, this Transformer enables classifying complex concepts, i.e., decoding neural data, with higher accuracy and with much less data by being pretrained in an unsupervised manner on a large corpus of unannotated neural recordings. Our approach generalizes to new subjects with electrodes in new positions and to unrelated tasks showing that the representations robustly disentangle the neural signal. Just like in NLP where one can study language by investigating what a language model learns, this approach opens the door to investigating the brain by what a model of the brain learns. As a first step along this path, we demonstrate a new analysis of the intrinsic dimensionality of the computations in different areas of the brain. To construct these representations, we combine a technique for producing super-resolution spectrograms of neural data with an approach designed for generating contextual representations of audio by masking. In the future, far more concepts will be decodable from neural recordings by using representation learning, potentially unlocking the brain like language models unlocked language.
Ignacio Cases, Christopher Wang, V. Subramaniam, A. Yaari, Gabriel Kreiman, B. Katz, Andrei Barbu · 7 authors totalDecentralized Learning Made Practical with Client Sampling
arXiv (Cornell University) · DOI 10.48550/arxiv.2302.13837 · 1 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Martijn de Vos, Akash Dhasade, Anne-Marie Kermarrec, Erick Lavoie, Sharma, Rishi · 6 authors totalAlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving
OSDI 2023 · arXiv 2302.11665 · Source: arxiv+semantic-scholarModel parallelism is conventionally viewed as a method to scale a single large deep learning model beyond the memory limits of a single device. In this paper, we demonstrate that model parallelism can be additionally used for the statistical multiplexing of multiple devices when serving multiple models, even when a single model can fit into a single device. Our work reveals a fundamental trade-off between the overhead introduced by model parallelism and the opportunity to exploit statistical multiplexing to reduce serving latency in the presence of bursty workloads. We explore the new trade-off space and present a novel serving system, AlpaServe, that determines an efficient strategy for placing and parallelizing collections of large deep learning models across a distributed cluster. Evaluation results on production workloads show that AlpaServe can process requests at up to 10x higher rates or 6x more burstiness while staying within latency constraints for more than 99% of requests.
Zhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu, Ying Sheng, Xin Jin, Yanping Huang, Zhifeng Chen · 11 authors totalAn Efficient B-tree Implementation for Memory-Constrained Embedded Systems.
CoRR · DOI 10.48550/arxiv.2302.07800 · arXiv 2302.07800 · Source: dblp+ubc-authorityRamon Lawrence, Nadir Ould-Khessal, Scott Fazackerley · 3 authors totalA Case Study on Record Matching of Individuals in Historical Archives of Indigenous Databases.
CoRR · DOI 10.48550/arxiv.2302.07784 · arXiv 2302.07784 · Source: dblp+ubc-authorityRamon Lawrence, Matthew Currie · 2 authors totalThe Capacity for Moral Self-Correction in Large Language Models
arXiv.org · DOI 10.48550/arXiv.2302.07459 · arXiv 2302.07459 · 217 citations · Source: semantic-scholarWe test the hypothesis that language models trained with reinforcement learning from human feedback (RLHF) have the capability to"morally self-correct"-- to avoid producing harmful outputs -- if instructed to do so. We find strong evidence in support of this hypothesis across three different experiments, each of which reveal different facets of moral self-correction. We find that the capability for moral self-correction emerges at 22B model parameters, and typically improves with increasing model size and RLHF training. We believe that at this level of scale, language models obtain two capabilities that they can use for moral self-correction: (1) they can follow instructions and (2) they can learn complex normative concepts of harm like stereotyping, bias, and discrimination. As such, they can follow instructions to avoid certain kinds of morally harmful outputs. We believe our results are cause for cautious optimism regarding the ability to train language models to abide by ethical principles.
Tom Brown, Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas Liao, Kamil.e Lukovsiut.e, Anna Chen, Anna Goldie · 48 authors totalScore-based Diffusion Models in Function Space
Journal of machine learning research · arXiv 2302.07400 · 93 citations · Source: semantic-scholarDiffusion models have recently emerged as a powerful framework for generative modeling. They consist of a forward process that perturbs input data with Gaussian white noise and a reverse process that learns a score function to generate samples by denoising. Despite their tremendous success, they are mostly formulated on finite-dimensional spaces, e.g., Euclidean, limiting their applications to many domains where the data has a functional form, such as in scientific computing and 3D geometric data analysis. This work introduces a mathematically rigorous framework called Denoising Diffusion Operators (DDOs) for training diffusion models in function space. In DDOs, the forward process perturbs input functions gradually using a Gaussian process. The generative process is formulated by a function-valued annealed Langevin dynamic. Our approach requires an appropriate notion of the score for the perturbed data distribution, which we obtain by generalizing denoising score matching to function spaces that can be infinite-dimensional. We show that the corresponding discretized algorithm generates accurate samples at a fixed cost independent of the data resolution. We theoretically and numerically verify the applicability of our approach on a set of function-valued problems, including generating solutions to the Navier-Stokes equation viewed as the push-forward distribution of forcings from a Gaussian Random Field (GRF), as well as volcano InSAR and MNIST-SDF.
Jean Kossaifi, Jae Hyun Lim, Nikola B. Kovachki, R. Baptista, Christopher Beckham, K. Azizzadenesheli, Vikram S. Voleti, Jiaming Song · 13 authors totalMachine Learning Model Attribution Challenge
First IEEE Conference on Secure and Trustworthy Machine Learning, Competition Track · arXiv 2302.06716 · Source: arxiv+ieee-satml+career-authorityDeepesh Chaudhari, Elizabeth M. Merkhofer, Hyrum S. Anderson, Keith Manville, Lily Wong, João Gante · 6 authors totalEstimation of Average Annual Daily Bicycle Count Using Bike-Share GPS Data and Bike Counter Data for an Urban Active Transportation Network.
CoRR · DOI 10.48550/arxiv.2302.06715 · arXiv 2302.06715 · Source: dblp+ubc-authorityRamon Lawrence, Marzi Rafieenia, Liza Wood, Mohsen Zardadi, Scott Fazackerley · 5 authors totalAn Application of Deep Learning for Sweet Cherry Phenotyping using YOLO Object Detection.
CoRR · DOI 10.48550/arxiv.2302.06698 · arXiv 2302.06698 · Source: dblp+ubc-authorityRamon Lawrence, Ritayu Nagpal, Sam Long, Shahid Jahagirdar, Weiwei Liu, Scott Fazackerley, Amritpal Singh · 7 authors totalData Selection for Language Models via Importance Resampling.
CoRR · DOI 10.48550/arXiv.2302.03169 · arXiv 2302.03169 · Source: dblp+stanford-authorityTengyu Ma, Sang Michael Xie, Shibani Santurkar, Tengyu Ma 0001, Percy Liang · 5 authors totalUsing Learned Indexes to Improve Time Series Indexing Performance on Embedded Sensor Devices.
CoRR · DOI 10.48550/arxiv.2302.03085 · arXiv 2302.03085 · Source: dblp+ubc-authorityRamon Lawrence, David Ding, Ivan Carvalho · 3 authors totalWitgenstein's influence on artificial intelligence
arXiv · DOI 10.48550/arxiv.2302.01570 · arXiv 2302.01570 · 0 citations · Source: openalexWe examine how much of the contemporary progress in artificial intelligence (and, specifically, in natural language processing), can be, more or less directly, traced back to the seminal work and ideas of the Austrian-British philosopher Ludwig Wittgenstein, with particular focus on his late views. Discussing Wittgenstein's original theses will give us the chance to survey the state of artificial intelligence, and comment on both its strengths and weaknesses. A similar text appeared first in Spanish as a chapter of CENTENARIO DEL SILENCIO (2021), a book celebrating 100 years since the publication of the Tractatus.
Piero Molino, Jacopo Tagliabue · 2 authors total