Papers.
Research connected to its authors, projects, companies, talks, events, and the rest of the graph.
Add a paper ↗System of Intelligent Actors: The DevOps chapter
0 citations · Source: semantic-scholarSunil Mallya, R. Manjunatha, Tatsuya Arai, Yinxiao Zhang, D. Rastogi, Goutham Nareddy, Nate Slater, Anshuman Mishra · 17 authors totalRisk and Russia's Takeover of Yukos: Why Russian and World Oil Production May Peak Even if There is No Scarcity of Oil
Fueling the Future: Prices, Productivity, Policies and Prophecies (USAEE/IAEE conference volume) · 0 citations · Source: google-scholarMarek Kolodziej, Douglas B. Reynolds · 2 authors totalHarpocrates: Oblivious Privacy in a Statically Typed World
0 citations · Source: semantic-scholarSinan Pehlivanoglu, Malte Schwarzkopf · 2 authors totalCluster Formation and Encrypted search in Big Data
2014 ASEE Zone 1 Conference Proceedings · DOI 10.18260/1-2-1153-54020 · 0 citations · Source: semantic-scholarGautam Siwach, Amir Esmailpour · 2 authors totalNSF CI Compass Virtual Workshop Report: AI meets CI - Intelligent Infrastructure for Major & Midscale Facilities
Zenodo (NSF CI Compass workshop report) · DOI 10.5281/zenodo.21383345 · 0 citations · Source: zenodo+openalexDean Wampler, Ewa Deelman, Charles Vardeman II, Prasanna Balaprakash, Gordon Broderick, Donald Brower, David Butcher, Kyle Chard · 35 authors totalShape Prior Fusion for Efficient 3D Reconstruction
TechRxiv (preprint) · DOI 10.36227/techrxiv.174285842.28885980/v2 · 0 citations · Source: crossrefCreating accurate 3D models from images is a challenging problem in computer vision, especially for applications like autonomous driving, virtual reality, and mobile computing. Traditional Multi-View Stereo (MVS) methods rely on geometry to reconstruct shapes but often struggle with missing details due to occlusions or textureless surfaces. In this work, we introduce a learning-based approach that enhances MVS by integrating deep priors, allowing for more complete and accurate 3D reconstructions with fewer input images. Our method refines fine details using masked silhouette and depth map losses, ensuring better shape recovery. Tested on the ShapeNet dataset, our model achieves a Chamfer Distance of 0.018 and an IoU of 0.80, outperforming classical MVS while also improving surface coverage to 91.1% and density score to 85.7%. By combining deep learning with traditional geometry, our approach makes 3D reconstruction more reliable and efficient, opening new possibilities for real-world applications.
Vincent Koc, Vamsidhar R. Kamanuru, Vinay Venkatesh, Hrishikesh Tawade, Sai Charan Dekkata · 5 authors totalWhat do news readers want?
NBER Working Paper 35289 · DOI 10.3386/w35289 · 0 citations · Source: semantic-scholarUsing a novel dataset covering the complete history of individual-level web traffic and digital subscriptions from a major metropolitan newspaper in the United States between 2020 and 2024, we investigate consumers' willingness to pay for different categories of news content, with particular focus on the kinds of coverage believed to generate civic externalities. Our identification strategy relies on the quasi-random arrival of paywall events which force consumers to subscribe if they wish to continue reading. Using this variation, we estimate a model of consumer demand and construct the optimal staff allocation for the paper under different counterfactual revenue models: a fully subscription-based model and a fully ad-supported model. Our results suggest that readers are willing to pay for local reporting, and that measures of demand based only on time-use substantially underestimate the value of “hard” news coverage on topics like local politics and public health. However, digital subscription revenues alone are insufficient to cover staff costs even at the highest revenue-generating sections of the paper. We use our model to estimate the subsidy required to expand the newspaper's production of investigative coverage.Institutional subscribers to the NBER working paper series, and residents of developing countries may download this paper without additional charge at www.nber.org.
Cameron Pfiffer, Gregory E. Martin, S. Vasserman · 3 authors totalA Common Language for Responsible AI: Methods and Insights from a Fieldwide Consensus-Building Effort
SSRN Electronic Journal · DOI 10.2139/ssrn.7273060 · 0 citations · Source: openalexDavid Talby, Matthew Elmore, Megan Salwei, Merage Ghane, Lisa Soleymani Lehmann, Shauna Overgaard, Naomi Lefkovitz, Cora Han · 24 authors totalINSURE-Dial: A Phase-Aware Conversational Dataset & Benchmark for Compliance Verification and Phase Detection
EACL 2026 · DOI 10.18653/v1/2026.eacl-long.237 · arXiv 2602.18448 · 0 citations · Source: semantic-scholar+arxivAdministrative phone tasks drain roughly 1 trillion USD annually from U.S. healthcare, with over 500 million insurance-benefit verification calls manually handled in 2024. We introduce INSURE-Dial, to our knowledge the first public benchmark for developing and assessing compliance-aware voice agents for phase-aware call auditing with span-based compliance verification. The corpus includes 50 de-identified, AI-initiated calls with live insurance representatives (mean 71 turns/call) and 1,000 synthetically generated calls that mirror the same workflow. All calls are annotated with a phase-structured JSON schema covering IVR navigation, patient identification, coverage status, medication checks (up to two drugs), and agent identification (CRN), and each phase is labeled for Information and Procedural compliance under explicit ask/answer logic. We define two novel evaluation tasks: (1) Phase Boundary Detection (span segmentation under phase-specific acceptance rules) and (2) Compliance Verification (IC/PC decisions given fixed spans). Per-phase scores are strong across small, low-latency baselines, but end-to-end reliability is constrained by span-boundary errors. On real calls, full-call exact segmentation is low, showing a gap between conversational fluency and audit-grade evidence.
Shiva Chaitanya, Shubham P. Kulkarni, Alexander Lyzhov, Preetam Joshi, S. Chaitanya · 5 authors totalStageleft: Multi-stage Programming in Standard Rust
International Conference on Generative Programming: Concepts and Experiences · DOI 10.1145/3814885.3816414 · 0 citations · Source: semantic-scholarRust has emerged as a popular systems language with growing interest in metaprogramming, yet it lacks staging support—developers must write unsafe, untyped macros instead. We present Stageleft, a library that brings type-safe staged programming to standard Rust without compiler modifications. Stageleft ensures hygienic code generation through ahead-of-time AST analysis, and handles free variables via a trait system that respects Rust's ownership rules. Stageleft demonstrates that staging can be practical and safe in Rust, enabling domain-specific optimizations while maintaining familiar developer interfaces.
Shadaj Laddad, Mingwei Samuel, Joseph M. Hellerstein · 3 authors totalIn Code They Think; In Proof We Trust
Queue · DOI 10.1145/3806226 · 0 citations · Source: openalex+semantic-scholarA preemptive strike against exfiltration
Erik Meijer · 1 author totalBeyond Semantic Similarity: Explicit Intent Modeling for Query-Product Matching
ACM conference proceedings · DOI 10.1145/3805712.3808500 · 0 citations · Source: openalex+orcidBuyer intent in e-commerce is multi-faceted and is expressed through explicit attributes—such as brand, size, color, and material, rather than through general topical relevance. However, many state-of-the-art scalable query-product matching systems rely on aggregate representations, scoring a single query embedding against a single item embedding. While efficient, this aggregation frequently fails to satisfy individual attribute intent: items can be semantically related, yet violate key aspects specified in the query. In contrast, fine-grained interaction methods can better capture aspect-level constraints, but are typically too expensive due to increased run-time computation and storage costs. We propose an aspect-aware ranking framework that retrieves and resolves aspects in queries and performs fine-grained semantic affinity match against aspects in products to compute an aggregate query-product level aspect affinity score. The proposed approach integrates (i) query aspect resolution (canonicalization) using structured aspect data, (ii) a model to learn granular aspect affinity signal capturing individual aspect-level understanding; and iii) an efficient design for online serving, significantly cutting cost associated with inference speed and storage. This design preserves the scalability of two-tower retrieval while substantially improving explicit intent satisfaction.
Alex Cozzi, Amanuel Alambo, Sathappan Muthiah, Diego Sierra, Zhenzhong Zhang, Atiq Islam · 6 authors totalNexa: Automatically Surfacing Business Impacting Insights in E-commerce Applications
CAIS '26: ACM Conference on AI and Agentic Systems · DOI 10.1145/3786335.3813185 · 0 citations · Source: dblp+crossref+semantic-scholarInternet-scale e-commerce storefronts serve millions of users (and increasingly user appointed agents). These storefronts are being rearchitected as compound AI systems with agentic workflows for customer interactions and backend processing. As this AI transformation and agentic economy is underway, product teams need to get actionable insights into business-impacting outcomes. Classical approaches such as static funnels or static dashboards cannot deal with the scale, diversity, and contextual interactions that happen over billions of user interactions. As such, we need novel agentic approaches to automatically surface business-impacting insights. We present Nexa, an agentic framework that surfaces business insights automatically. We formalize the target of automated insight discovery in terms of Contrastive Stateful Trajectories (CST): a structural specification over contextual and sequential behavioral patterns whose presence or absence significantly shifts a business KPI across user cohorts. Nexa satisfies three design requirements simultaneously: expressivity through the CST abstraction, scalability through a custom analytics backend for CST computations, and explainability by overlaying usable presentation layers for analysts to verify the insights. We demonstrate Nexa on representative workloads and show that it surfaces actionable contextual patterns spanning user, app, agent, and backend behaviors.
Evan Chan, Smart Sun, Sayan Sinha, Haijie Wu, Joel Goldfoot, Aditya Ganjam, Jibin Jhan, Qichu Gong · 17 authors totalTracking Capabilities for Safer Agents
CAIS '26: ACM Conference on AI and Agentic Systems · DOI 10.1145/3786335.3813127 · arXiv 2603.00991 · 3 citations · Source: arxivAI agents that interact with the real world through tool calls pose fundamental safety challenges: agents might leak private information, cause unintended side effects, or be manipulated through prompt injection. To address these challenges, we propose to put the agent in a programming-language-based "safety harness": instead of calling tools directly, agents express their intentions as code in a capability-safe language: Scala 3 with capture checking. Capabilities are program variables that regulate access to effects and resources of interest. Scala's type system tracks capabilities statically, providing fine-grained control over what an agent can do. In particular, it enables local purity, the ability to enforce that sub-computations are side-effect-free, preventing information leakage when agents process classified data. We demonstrate that extensible agent safety harnesses can be built by leveraging a strong type system with tracked capabilities. Our experiments show that agents can generate capability-safe code with no significant loss in task performance, while the type system reliably prevents unsafe behaviors such as information leakage and malicious side effects.
Martin Odersky, Yaoyu Zhao, Yichen Xu, Oliver Bračevac, Cao Nguyen Pham · 5 authors totalWhose Knowledge Counts? Co-Designing Community-Centered AI Auditing Tools with Educators in Hawai'i
CHI · DOI 10.1145/3772318.3790958 · arXiv 2603.16646 · 1 citations · Source: arxiv+semantic-scholarAlthough generative AI is being deployed into classrooms with promises of aiding teachers, educators caution that these tools can have unintended pedagogical repercussions, including cultural misrepresentation and bias. These concerns are heightened in low-resource language and Indigenous education settings, where AI systems frequently underperform. We investigate these challenges in Hawai`i, where public schools operate under a statewide mandate to integrate Hawaiian language and culture into education. Through four co-design workshops with 22 public school educators, we surfaced concerns about using generative AI in educational settings, particularly around cultural misrepresentation, and corresponding designs for auditing tools that address these issues. We find that educators envision tools grounded in specific Hawaiian cultural values and practices, such as tracing the genealogy of knowledge in source materials. Building on these insights, we conceptualize AI auditing as a community-oriented process rather than the work of isolated individuals, and discuss implications for designing auditing tools.
Michael Ryan, Dora Zhao, Hannah Cha, Michael J. Ryan, Angelina Wang, Rachel Baker-Ramos Evyn-Bree Helekahi-Kaiwi, Rebecca Diego, Josiah Hester · 8 authors totalTimelyLLM: Time-sensitive LLM Serving System for Physical-I/O Limited Agents
ACM SIGMOBILE International Conference on Mobile Systems, Applications, and Services · DOI 10.1145/3745756.3809203 · 0 citations · Source: semantic-scholarLarge Language Models (LLMs) are increasingly integrated into Physical-I/O limited agents, such as robots and voice assistants, which execute outputs sequentially. However, existing LLM serving systems typically employ a throughput-oriented batching mechanism, ignoring the large gap between LLM generation speed and the constrained physical I/O rates of agents, thus wasting execution slack and worsening resource contention. Besides, they treat all tokens equally and cannot anticipate the execution implications of different content, preventing scheduling aligned with agent-side behavior. To address it, we propose a new system named TimelyLLM that coordinates LLM generation with the physical behavior of agents. TimelyLLM introduces a novel segmented generation and scheduling mechanism, strategically leveraging the time gap between agent plan generation and execution to reduce contention and improve response latency under multi-agent workloads. We implement TimelyLLM on top of a widely-used LLM serving framework. We also build a dataset collection system to construct serving workloads from real-world robots, including drones, robot arms, and quadruped robots. Our evaluation demonstrates that TimelyLLM improves the time utility up to 1.52×, and reduces the overall waiting time by 84%.
Anurag Khandelwal, Neiwen Ling, Guojun Chen, Lin Zhong · 4 authors totalBehavioral Transfer via Automated Prompt Optimization and LLM-as-a-Judge Evaluation Loops for Prompt-Based Knowledge Distillation
2026 Systems and Information Engineering Design Symposium (SIEDS) · DOI 10.1109/sieds69358.2026.11540303 · 0 citations · Source: crossrefLarge language models (LLMs) are efficient and expensive to implement. We present APO-KD, a framework that uses zero-fine-tuning to replicate the observable behavior of a teacher LLM using a cheaper student by maximising discrete packages of prompts rather than weights. APO-KD (i) produces teacher reference outputs, (ii) executes the student with candidate prompts, (iii) assesses alignment with an LLM-as-a-Judge rubric in terms of answer quality, format fidelity, constraint adherence, and consistency, and (iv) rewrites prompts based on judge feedback. On limited bullet-point summarization and code generation (function and unit tests only), distilled prompts are much more effective in getting students to comply and lowering the behavior gap to the teacher compared to zero-shot and manual prompts, and can be deployed quickly and with control using a small budget when fine-tuning is not feasible.
Vincent Koc · 1 author totalAgent-Centric Column-Aware Prompt Optimization for Structured Data Tasks with LLMs
2026 International Seminar on Intelligent Business and Edge-Computing Research (ISIBER) · DOI 10.1109/isiber68248.2026.11470127 · 0 citations · Source: crossrefLarge language models (LLMs) are being used more commonly in applications that use structured data (tables and databases) in reasoning. Nonetheless, naive few-shot prompt selection tends to disregard the underlying schema, which results in poor performance and higher rates of hallucinations. The paper presents a column-aware few-shot prompt optimization model that uses schema knowledge and a multi-objective search based on Optuna to sample and rank examples to solve structured data problems. The proposed system has tunable parameters, which include the selection of columns, examples, and formatting, and are optimized based on the desired metrics, including the accuracy of the tasks and the rate of hallucinations. Table question answering, report generation and analytics assistant task experiments show that faster convergence and better performance are realized by column-aware optimization than by random or best-of-one/best-of-three selection. The framework is also incorporated into an agent optimization platform, which allows automatic and adaptive timely management of production environments.
Vincent Koc, Vamsidhar R. Kamanuru · 2 authors totalBulletTime: Time Dilation for High-Fidelity Tracing
International Symposium on Computer Architecture · DOI 10.1109/ISCA66397.2026.00125 · 0 citations · Source: semantic-scholarMuch of computer systems and architecture research depends on accurate, high-fidelity program tracing for simulation, profiling, and debugging. Unfortunately, with significant improvements in compute and memory instruction execution speeds, tracing frameworks incur frequent I/O to persist traced events to disk. We find that such delays can be bursty and asymmetric across application threads, resulting in an inadvertent reordering of application and system operations relative to untraced execution. In our analysis, such reordering often leads to significant changes in the behavior of the studied application, thereby contaminating insights from simulation and profiling studies of the corresponding captured traces. In this work, we formalize the application behavior under study to establish correctness requirements for traced application execution in the presence of tracing-induced delays. We propose a novel time dilation approach that strategically slows execution for application and system threads while meeting correctness requirements. We implement the time-dilation approach in Bul-letTime, a tracing framework built atop Pin, the de facto binary instrumentation tool. We evaluate BulletTime for memorycontiguity and synchronization studies on real-world applications and workloads. Our results show that while existing tracing approaches can cause application behavior to deviate by as much as $20 \times$ compared to untraced execution, BulletTime's deviations are < 10% even in extreme cases of asymmetric tracing delays.
Anurag Khandelwal, Michael Wu, Sibren Isaacman, Abhishek Bhattacharjee · 4 authors totalHigh-Performance Durable LLM Observability at Scale Architecture and Benchmarking of the Opik Platform
2026 IEEE 15th International Conference on Communication Systems and Network Technologies (CSNT) · DOI 10.1109/csnt69054.2026.11502391 · 0 citations · Source: crossrefThe current trend to implement large language models (LLMs) in production settings has generated a must-have requirement to embed scalable, lasting, and high-fidelity observability solutions. Conventional logging and ephemeral tracing is not detailed enough to capture the interactions which are complex and multi-turned as well as tool augmented of the modern LLM applications. This paper describes Opik, a high-performance observability platform, developed to provide both long-term and real-time trace retention of systems based on LLM. Opik has an asynchronous ingestion pipeline, hierarchical trace models, and a hybrid storage design that combines time-series and a vector database. By stringent benchmarking against the set standards we prove that Opik has better ingestion latency, efficient storage and query performance, enabling advanced analytics and continuous regression analysis. The outcomes of our work set Opik to be a future-proof implementation of the LLM observability at enterprise level.
Vincent Koc, Sagheer Abbas · 2 authors totalFairness and Bias Management in Health AI: consensus-Based Recommendations for Best Practices Across the AI Lifecycle
The American Journal of Bioethics · DOI 10.1080/15265161.2026.2677391 · 0 citations · Source: openalexIn a landscape marked by uneven oversight, consensus-based guidance plays a crucial role in advancing AI fairness and managing bias. It contributes to a shared understanding across disciplines, ena...
David Talby, Merage Ghane, Matthew Elmore, Irene Dankwa‐Mullan, Shawn Stapleton, Kellie Owens, Sana Khalid, Allie Delonay · 14 authors totalImpact of using artificial intelligence as a second reader in breast screening including arbitration
Nature Cancer · DOI 10.1038/s43018-026-01128-z · 5 citations · Source: semantic-scholarThe impact of incorporating artificial intelligence (AI) into a double-read breast-screening workflow, including arbitration, is unclear. This retrospective study included 50,000 representative women from two NHS breast-screening centers. All the women had long-term follow-up, allowing us to determine whether use of AI leads to earlier cancer detection. Cases requiring arbitration (8,732 cases) were read by 22 readers in a reader study, following their normal arbitration workflow. Overall, after arbitration, replacing the second reader with AI was noninferior (5% margin) to two human readers in terms of sensitivity and specificity (P < 0.001) while offering a workload benefit. Arbitration improved the specificity of the AI arm by overruling cases incorrectly recalled by the AI tool; however, it also overruled the AI tool recall decision for some interval and next-round cancers. Further development of the AI tool alongside improvement in its explainability could lead to the earlier detection of cancers.
Daniel Golden, L. Warren, J. Venton, Kenneth C Young, M. Halling-Brown, Christopher J. Kelly, Marc Wilson, Megumi Morigami · 34 authors totalDiagnostic accuracy, fairness and clinical implementation of AI for breast cancer screening: results of multicenter retrospective and prospective technical feasibility studies
Nature Cancer · DOI 10.1038/s43018-026-01127-0 · 5 citations · Source: semantic-scholarArtificial intelligence (AI) promises to enhance breast cancer screening. Here we evaluated Google’s mammography AI system (version 1.2) across two phases: a retrospective study using 115,973 mammograms from five National Health Service screening services with 39-month follow-up and prospective noninterventional feasibility deployment at 12 sites (9,266 cases). The primary endpoint was AI sensitivity and specificity versus first reader using a 5% noninferiority margin. The secondary endpoints were performance versus second or consensus readers and breast-level analyses. Retrospectively, AI achieved superior sensitivity (0.541 versus 0.437 for first reader, P < 0.001) and noninferior specificity (0.943 versus 0.952, P < 0.001). Cancer detection rate increased from 7.54 to 9.33 per 1,000 women, with AI detecting 25.0% of interval cancers. Performance was particularly strong for first screens (39.3% fewer recalls, 8.8% higher detection) and invasive cancers. No systematic demographic disparities were observed. Simulated second-reader replacement reduced reading time by 32% while increasing detection by 17.7%. Prospective deployment confirmed technical feasibility but revealed a distribution shift requiring threshold recalibration. Implementation requires adaptive calibration and continuous monitoring to ensure safety and equity.
Daniel Golden, Christopher J. Kelly, Marc Wilson, L. Warren, Richard Sidebottom, M. Halling-Brown, Lin Yang, Megumi Morigami · 38 authors totalSmartCert: A Multi-modal framework for automated guided vehicle screening
Pervasive and Mobile Computing · DOI 10.1016/j.pmcj.2025.102127 · 0 citations · Source: crossrefVincent Koc, Xu Chen, Sandeep Kanta, Santhi Bharath Punati, Arif Hussain, Sunny Katyara · 6 authors totalAssessing the robustness of evaluation metrics for synthetic ECG signal quality
Computers in Biology and Medicine · DOI 10.1016/j.compbiomed.2026.111824 · 0 citations · Source: orcidGonçalo Martins Ribeiro, Maria Russo, Inês Sousa, Ricardo Santos, Joana Rebelo, Fabiana Clemente, Gonçalo Ribeiro, André Carreiro · 8 authors totalPlanning the development of an AI-driven decision support architecture for the recognition of sudden cardiac arrest by 9-1-1 telecommunicators: report of a community engagement and brainstorming meeting.
CJEM · DOI 10.1007/s43678-026-01232-0 · 0 citations · Source: semantic-scholarRandy Giffen, Christian Vaillancourt, S. Leduc, Sarika Naidoo, M. Charette, J. Phillip Nicholson, M. Church, Wojtek Michalowshi · 18 authors totalGeneralization of AI-Based Gestational Age Assessment Using Blind Sweep Ultrasonography
JAMA Network Open · DOI 10.1001/jamanetworkopen.2026.22484 · 1 citations · Source: semantic-scholarKey Points Question Can artificial intelligence (AI) models trained on blind sweep ultrasonography scans generalize to new clinical settings and perform as well as traditional sonographers on estimating gestational age? Findings In this diagnostic study of 385 participants, the AI model effectively generalized to new clinical environments and institutions, achieving a mean absolute error of 4.2 days, which was noninferior to the clinical standard. Meaning This AI system demonstrated strong potential to expand access to diagnostic tasks such as gestational age estimation, particularly in low-resource settings, by enabling novice operators to perform accurate ultrasonography assessments.
Daniel Golden, Angelica Willis, Chace Lee, Justin D. Krogue, A. Wickramanayake, Nichole Young-Lin, Stacey Caron, Priscah Cheruiyot · 28 authors totalModel Card for OpenAI Privacy Filter
arXiv 2608.18274 · 0 citations · Source: arxivOpenAI Privacy Filter is a compact, bidirectional token-classification model for detecting and redacting personally identifiable information (PII) and secrets in unstructured text. The model is derived from an autoregressively pretrained checkpoint and converted into a bidirectional, banded-attention classifier that labels an input sequence in a single forward pass. A constrained Viterbi decoder produces coherent spans across eight privacy categories and exposes configurable operating points for precision-recall tradeoffs. Privacy Filter has 1.5 billion total parameters, 50 million active parameters per token, and a 128,000-token context window. It is designed for efficient local deployment and domain-specific fine-tuning. Privacy Filter is intended as a configurable data-minimization component within layered privacy workflows, not as an anonymization or compliance guarantee.
Mihai Maruseac, Charles de Bourcy, Sahra Ghalebikesabi, Avi Schwarzschild, Alex Gorbachev, Annie Chu, V. Kyrylov, Tong Mu · 25 authors totalEigenius: A Typed Knowledge-Graph DBMS with Epistemic Stratification and Institution-Mediated Reasoning
arXiv · arXiv 2608.04457 · 0 citations · Source: arxivMatthew Fuchs, Hans-Martin Will, Allen L. Brown · 3 authors totalAI Security Priorities: A Field-Wide Agenda
arXiv preprint · DOI 10.48550/arXiv.2607.26069 · arXiv 2607.26069 · 0 citations · Source: arxiv+dblpAs AI systems are rapidly integrated into critical economic, governmental, and national security functions, the gap between AI adoption and AI security readiness continues to widen. This paper presents a prioritized agenda for advancing AI security, informed by structured interviews with leaders across industry, government, and civil society, and refined through a multi-sector expert workshop. Participants identified and ranked the highest-importance and most cost-effective areas where progress could strengthen AI security - from protecting frontier AI systems and their underlying infrastructure to improving cybersecurity practices as AI reshapes the threat landscape. The resulting priorities are organized across four themes: establishing strategic foundations and policy frameworks; advancing public-private coordination and institutional infrastructure; advancing technical security engineering and assurance; and governing agentic AI under adversarial pressure. For each priority area, expert authors provide detailed analyses that define the problem, assess the current landscape, and identify actionable projects that stakeholders across sectors can pursue. The paper aims to serve as an initial practical foundation for coordinated investment and action across the AI security field. It is designed to serve both current practitioners and individuals and organizations looking to enter the field by identifying concrete, high-impact contributions suited to a range of strengths and capacities.
Buck Shlegeris, Gil Gekker, Rachel Steratore, Everett Smith, Asher Brass-Gershovich, Varun Gandhi, Nicole Nichols, Vijay Bolina · 12 authors totalClassifying Capabilities (Extended Version)
arXiv 2607.24504 · 0 citations · Source: arxivCapture checking in Scala 3 enables lightweight and practical effect and resource tracking by recording capabilities in types. However, the system offers no way to reason about kinds of capabilities. Natural constraints such as "retaining only the control-flow capabilities of this closure" or "excluding all thread-local capabilities from this argument" become inexpressible. Both arise in the Scala 3 standard library: "Try" re-throws caught exceptions, so it retains only the control-flow capabilities of its body, and "Future" must not capture thread-local resources. The inability to state these constraints has kept parts of the library outside capture checking. We introduce capability classifiers: a tree-structured, user-extensible hierarchy of tags that classify capabilities by their semantic role. Projections filter capture sets by classifier, supporting both inclusion ("c.only[C]") and exclusion ("c.except[C]"). The tree structure enables decidable disjointness reasoning: classifiers on separate branches are guaranteed to be disjoint regardless of unknown extensions elsewhere in the hierarchy. We formalize classifiers as an extension of System Capless, a core calculus for capture checking, introducing a classifier kind algebra based on intersection, union, and subtraction of classifier subtrees. We extend the operational semantics to model exception interception and establish type safety, effect safety, and handler coverage via a big-step proof, fully mechanized in Lean 4. Classifiers are implemented in the Scala 3 capture checker, and we demonstrate their use on standard library types and real-world effect exclusion patterns.
Martin Odersky, Cao Nguyen Pham, Oliver Bračevac, Yichen Xu, Yaoyu Zhao · 5 authors totalLeveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection
arXiv · arXiv 2607.23961 · 0 citations · Source: semantic-scholarRecent advances in speech synthesis and voice conversion, which pose threats to security and privacy, have underscored the need for deepfake detection technology. Although existing detection systems achieve strong performance on individual datasets, they often fail to generalize across diverse datasets. Prior methods for improving generalization, including data augmentation, adversarial training on auxiliary factors such as language or codec types, and Mixture-of-Experts (MoE), are limited by predefined augmentation coverage, difficulties in obtaining auxiliary factors, and substantial model complexity. In this work, we propose a practical dataset-aware framework for deepfake detection. Our method targets heterogeneous datasets for which auxiliary annotations such as language, codec, or spoofing method may not be consistently available. We therefore rely only on dataset identity as a naturally available supervisory signal for multitask (MT) and gradient reversal layer (GRL) training, allowing the model to investigate both dataset-aware multitask supervision and adversarial suppression of dataset-specific information. We conduct experiments following the 2025 Speech DeepFake Arena benchmark protocol, evaluating our model across multiple evaluation datasets and reporting aggregate performance in terms of Equal Error Rate (EER), including Average EER and Pooled EER. Compared with the baseline, MT reduces Average EER by 13.14% relatively, while GRL reduces Pooled EER by 5.32% relatively. These results demonstrate that our method can improve aggregate detection performance across heterogeneous evaluation datasets, offering a practical solution for deploying reliable deepfake detection systems on diverse and unseen real-world data.
Yishay Carmiel, Mingrui Liang, Thomas Thebaud, Lukasz Wójciak, Laureano Moro Velázquez, Jesus Villalba Lopez, N. Dehak · 7 authors totalA Coulomb Particle Model for Learning Kernel Attention in Transformers
arXiv · DOI 10.48550/arXiv.2607.23869 · arXiv 2607.23869 · 0 citations · Source: openalex+arxivRandomized features provide a scalable approximation to kernel machines, but their performance depends strongly on the choice of feature distribution. We propose a particle-based method that learns this distribution by optimizing kernel-target alignment while regularizing particles with a Riesz/Coulomb repulsive potential. The resulting Hamiltonian yields diverse, task-adaptive random features and admits a mean-field description through a McKean--Vlasov equation. We instantiate the method in linearized Transformer attention by learning positive random-feature maps in a first alignment phase, then freezing the kernel and training the remaining network parameters with cross-entropy. Experiments on synthetic classification and sentence-level benchmarks show that learned kernelized attention can improve accuracy, calibration, and robustness for several feature maps while preserving linear-attention inference complexity.
Alex Cozzi, Masoud Badiei Khuzani, Sharath Honnaiah, Atiq Islam, Abraham Bagherjeiran · 5 authors totalDeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs
arXiv · arXiv 2607.22555 · 0 citations · Source: openalexMedical diagnosis is a multi-stage process: extract facts, consult knowledge, generate a differential analysis, and select the best diagnosis with explanations. Frontier LLMs are strong generalists, but single-shot prompting often yields brittle diagnostic reasoning. We present the DeepLens Diagnosis Agent, a five-stage harnessing pipeline (combining model capabilities with disciplined process constraints) centered on a small medical reasoning model (JSL Medical Small 7B v2) and retrieval-augmented generation (RAG). The pipeline enforces structured clinical extraction, disciplined retrieval, constrained candidate generation, explicit evidence triangulation, and an auditable final decision. On the 915-case DiagnosisArena benchmark, the agent achieved 60.14% top-1 diagnostic accuracy, the highest among small and medium-sized models. The same model without the agent workflow achieved 23.99%, a +36-point gain from workflow design alone, despite 88.2% on standard medical benchmarks, showing that diagnostic reasoning under uncertainty requires more than knowledge recall. The agent costs USD 0.0072 per case (24K tokens on A100) with 24-second latency, 35-45% cheaper than Claude Sonnet 4.5 (USD 0.0110) and Gemini 3.1 Pro (USD 0.0128) while outperforming them by +9.70pp and +9.17pp. Harnessing can also correct frontier model failures; workflow constraints can outweigh parameter count or API cost. Beyond aggregate accuracy, the pipeline produces structured intermediate artifacts that make each stage inspectable and support error localization. These properties support high-stakes settings where traceability, reproducibility, and auditable evidence matter alongside benchmark performance.
David Talby, Mahmood Bayeshi, Veysel Kocaman, Muhammed Ali Naqvi, Yigit Gul · 5 authors totalSystem Capybara: Tracking Capabilities for Separation and Freshness (Extended Version)
arXiv 2607.09383 · 0 citations · Source: arxivSubstructural type systems give strong static control over aliasing. Examples include uniqueness, separation, and borrowing. How can such control be brought to established languages whose programming models rely on higher-order abstraction, unrestricted aliasing, and pervasive sharing? We study this problem in the context of Scala. We show how to retrofit these guarantees selectively instead of globally: ordinary code keeps Scala's usual aliasing discipline, while stronger guarantees can be enforced where they matter. Our starting point is Scala's capture checking, whose treatment of capabilities is inspired by the object-capability tradition: capabilities are ordinary values, and capture sets record, in a value's type, which capabilities the value may use. We develop System Capybara, which adds a selective alias-control layer to this mechanism. By tracking separation, consumption, freshness, and read-only access for capabilities, Capybara recovers key reasoning principles from substructural and ownership-based disciplines without global invariants. We give a type-preserving translation from the surface calculus Capybara to CoreCapybara, a core calculus extending System Capless, the earlier foundation for capture checking. The translation uses quantifiers for capture polymorphism and freshness, and constraint-indexed modal types for separation. We prove a semantic soundness result for the core calculus in Lean 4 and derive type safety, memory safety (no use-after-free or double-free), immutability of read-only computations, and data-race freedom for well-typed programs. Finally, we implement Scala 3's new separation checker, which brings higher-order separation reasoning about effects, capabilities, and resources to ordinary Scala, including fearless concurrency.
Martin Odersky, Yichen Xu, Oliver Bračevac, Cao Nguyen Pham, Yaoyu Zhao · 5 authors totalGemma 4 Technical Report
arXiv (Google DeepMind technical report) · arXiv 2607.02770 · 41 citations · Source: semantic-scholar+arxivWe introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parameters. Alongside improved vision and audio encoders for all model sizes, we propose a unified, encoder-free architecture for our 12B model, which ingests raw audio and image patches. Furthermore, we integrate a thinking mode, enabling Gemma models to generate reasoning traces prior to responding. We improve inference speed, memory, and compute efficiency, as well as long-context abilities through critical design choices. Gemma 4 establishes a leap in performance across STEM, multimodal, and long-context benchmarks, and rivals larger, frontier open models in human-rated tasks.
Clément Farabet, Mayank Chaturvedi, Michelle Casbon, Gemma Team, Sherif El Abd, Vaibhav Aggarwal, Robin Algayres, Alek Andreev · 323 authors totalAuto-FL-Research: Agentic Search for Federated Learning Algorithms
arXiv preprint · arXiv 2607.01366 · 0 citations · Source: arxiv+semantic-scholarFederated learning (FL) research often depends on many small but consequential algorithmic choices: optimizer variants, server aggregation rules, local training schedules, normalization, regularization, and model architecture. These choices are expensive to explore manually and difficult to compare fairly when candidate changes can also alter the FL training or evaluation path. In this work, we present Auto-FL-Research (AFR), a constrained coding-agent workflow for FL algorithmic recipe search. Agents may propose and implement candidate training algorithms, including server aggregation rules, client update schedules, local objectives, and registered model variants, while task profiles fix the mutation surface, compute budget, communication contract, and final model evaluation. Each campaign records candidate scores, runtime, edited files, artifacts, and failure status. We evaluate AFR on five healthcare cross-silo FLamby tasks and on grouped-client profiles for the five fixed LEAF datasets plus the LEAF synthetic task. Five-seed repeat evaluations support gains on four FLamby tasks and five of six LEAF profiles, while also exposing seed-sensitive and search-selected failure cases. Same-budget controls show that several gains correspond to FL-recipe changes, whereas other improvements are recovered by fixed-surface scalar controls or fail under repeat or held-out evaluation. These mixed outcomes are part of the contribution: they show how agent-generated candidates can be separated into repeated FL mechanisms, fixed-surface tuning effects, and selected single-run artifacts.
Chester Chen, Holger R. Roth, Ziyue Xu, Daguang Xu, Peter Cnudde, Andrew Feng · 6 authors totalFearless Concurrency on the GPU
arXiv · DOI 10.48550/arXiv.2606.15991 · arXiv 2606.15991 · 0 citations · Source: semantic-scholar+dblpRust has made safe systems programming practical on the CPU, but writing custom GPU kernels in Rust still forces programmers outside the language's ownership guarantees. We present cuTile Rust, a tile-based system for safe, idiomatic GPU kernel authoring in Rust. cuTile Rust extends Rust's ownership discipline to tile-based GPU kernels: mutable outputs are split into disjoint pieces, kernel launches preserve the host-side ownership contract, and programmers can opt out locally when they need lower-level control. The system also provides a composable host execution model spanning synchronous launches, asynchronous pipelines, and CUDA graph replay. Our evaluation shows that these abstractions can preserve performance on high-end GPUs. On the NVIDIA B200 GPU, cuTile Rust achieves 7 TB/s for element-wise operations and 2 PFlop/s for GEMM (96% of cuBLAS), matching cuTile Python within measurement noise. Grout, a cuTile-Rust-based inference engine, exercises cuTile Rust across an end-to-end Qwen3 inference path. In batch-1 decode, Grout reaches 171 generated tokens/s for Qwen3-4B on the NVIDIA GeForce RTX 5090 and 82 generated tokens/s for Qwen3-32B on the B200, competitive with vLLM and SGLang and consistent with an HBM roofline sanity check.
Jared Roesch, Melih Elibol, Isaac Gelado, E. Buehler, Michael Garland · 5 authors totalBridging the Semantic-Collaborative Gap: An Asymmetric Graph Architecture for Cold-Start Recommendation
arXiv · arXiv 2606.06225 · 0 citations · Source: semantic-scholarCollaborative filtering and graph-based recommendation models are highly effective because they leverage observed user interactions, but this dependence creates a fundamental cold-start challenge when newly added content has no interaction history. In Tubi's production retrieval system, this challenge is further constrained by the serving interface: new content must be assigned a standalone embedding immediately, and the model must also produce device embeddings suitable for approximate nearest-neighbor retrieval. We address this setting by formulating cold-start recommendation as an inductive graph-completion problem on a temporal bipartite device-content graph. We propose Shallow-RHS, an asymmetric link-prediction architecture in which the left-hand side (LHS) device tower leverages temporally valid watch-history message passing to capture collaborative signals, while the right-hand side (RHS) content tower is intentionally shallow with respect to the graph and encodes content solely from intrinsic features. The RHS tower does not use ID-based embeddings, content-side subgraphs, neighbor aggregation, or interaction-derived representations, forcing the content encoder to map intrinsic features into a collaborative-filtering-aware embedding space. After training, the learned content encoder generates embeddings for both warm and newly ingested content, enabling implicit graph completion through retrieval of warm surrogate neighbors. We further extend the same representation-completion principle to device cold-start by constructing cohort-based embeddings from demographic features. Large-scale online experiments demonstrate consistent relative improvements in content cold-start engagement, promotion speed, impression acquisition, and device cold-start engagement.
Mike Tamir, Anh Truong, John Trenkle, Yuanbo Chen, Honghong Zhao, Abdullah Alchihabi, Eric Fang, Michael Tamir · 8 authors totalGate AI: LLM Security Benchmark Evaluation Methodology and Results
arXiv preprint (cs.LG, cs.CR) · DOI 10.48550/arXiv.2606.02959 · arXiv 2606.02959 · 0 citations · Source: arxivPublished evaluations of prompt-injection and jailbreak detectors for Large Language Models often suffer from two systematic weaknesses: per-dataset threshold tuning and undisclosed operating points. We describe an evaluation harness that addresses both. The detector under evaluation is scored across 16 public benchmarks (12,111 samples) using 5-fold cross-validation. StratifiedKFold (by row) is the headline pass; a parallel StratifiedGroupKFold pass over a composite key (parent-prompt id plus MinHash + LSH near-duplicate clusters at Jaccard >~ 0.8) runs alongside it as a leakage-premium diagnostic. A single global operating point is selected on the held-out folds (max F1 subject to FPR <= 1%) and applied uniformly to every dataset, so per-dataset results reflect one threshold rather than per-benchmark optimisation. Generalisation is examined through a battery of diagnostics (leave-one-dataset-out cross-validation, a random-label control, adversarial validation, permutation feature importance, length-bias correlation, classifier-head agreement, cross-source near-duplicate detection, threshold transferability, train-vs-OOF agreement, and a paraphrase-invariance probe), most with a quantitative pass threshold and the remainder with a stated failure mode. For every external comparison, the detector's threshold is re-tuned to the competitor's published false-positive rate so head-to-head values are evaluated at matched operating points.
Ryle Goehausen, Marcus Sousa · 2 authors totalHalf the Interference, Most of the Answer: Approximate Quantum Simulation via Path-Sum Pruning
arXiv · arXiv 2606.01922 · 0 citations · Source: semantic-scholarClassical simulation of quantum circuits is expensive for two distinct reasons. The obvious one is state-space size: an n-qubit system requires exponentially many amplitudes. The less obvious one is interference: useful output distributions emerge only after many computational histories have been coherently combined at common endpoints, and this aggregation step is itself a substantial source of cost. We introduce statistical interference sampling, a framework that makes this second bottleneck explicit by treating endpoint interference as a separately schedulable computation. Using the Chemical Abstract Machine (ChAM) as our model, weighted path contributions evolve as concurrent molecular species, and interference reactions combine contributions that share a common output state. A threshold rule terminates the process once an endpoint accumulates sufficient amplitude, discarding the remaining reactions. The method does not improve worst-case complexity and is not intended as a general-purpose simulator. Its purpose is to ask a more targeted question: how much of the interference calculation can be skipped while still recovering a useful output distribution? On benchmark circuits for Deutsch-Jozsa, Grover search, Simon's problem, and small Shor period-finding instances, we find that nearly 50% of endpoint interference reactions can be omitted while maintaining over 90% output accuracy for most algorithms tested. These results suggest that interference arithmetic is a structured resource that admits meaningful approximation, and that exposing it explicitly opens new opportunities for pruning strategies across path-sum, Pauli-path, and tensor-network simulation methods.
Sinan Pehlivanoglu, S. Iyengar, Amr Sabry · 3 authors totalClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree
arXiv · DOI 10.48550/arXiv.2606.01494 · arXiv 2606.01494 · 1 citations · Source: arxivAgent skills extend AI agents with reusable instructions, tools, scripts, references, and workflows, establishing a security boundary distinct from both model safety and traditional package-malware detection. ClawHub Security Signals is a sanitized dataset of 67,453 latest public OpenClaw skill versions. Each row pairs redacted SKILL.md content and sanitized bundled files where present with a final ClawScan registry verdict and evidence from three scanner families: VirusTotal, static heuristic analysis, and NVIDIA SkillSpector. Rather than estimating malicious-skill prevalence, we study scanner disagreement. The three scanners rarely flag the same skills: any pair overlaps on at most 10.4% of their combined positives, only 0.69% of skills are flagged by all three, and 81.9% of flagged skills are identified by a single scanner. The disagreement is structured by attack surface. SkillSpector, which raises semantic agentic-risk advisories rather than malware-reputation signals, is positive for 19,209 of 25,504 suspicious rows (75.3%) but only 14 of 206 malicious rows (6.8%). The malicious-verdict region shows the inverse profile: 150 of 206 malicious rows (72.8%) are VirusTotal-positive, consistent with bundled-code malware evidence. These results show that agent-skill security requires layered governance, not single-scanner allow/block decisions. The corpus is released as a sanitized silver-standard dataset: labels are the registry's automated verdicts, not human-annotated ground truth, and the release represents an early, versioned snapshot intended to support the community while a human-annotated subset is developed. Further research is encouraged, including models tailored for skill-security triage.
Vincent Koc, Patrick Erichsen, Jacob Tomlinson, Agustin Rivera, Michael Appel, Nir Paz · 6 authors totalA Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models
arXiv · arXiv 2606.00027 · 0 citations · Source: openalexLarge language models (LLMs) are increasingly deployed across healthcare, yet existing benchmarks fail to capture model behavior under adversarial or ethically complex conditions common in clinical practice. We developed a multi-domain red teaming framework evaluating eleven contemporary LLMs across 690 clinically grounded scenarios spanning nine domains and over 150 subcategories. Scenarios incorporated adversarial transformations, and responses were assessed using a seven-dimension rubric with LLM-assisted scoring and human-in-the-loop validation. Results revealed substantial performance variance, with mean scores ranging from 0.791 to 0.984. Critically, several high-performing systems produced complete failures in individual safety-critical scenarios, demonstrating that aggregate accuracy masks clinically meaningful risk. The highest-performing systems (X-BAI, GPT-5, Claude Opus 4.1) achieved scores above 0.97 with low variance, while performance varied significantly across domains. Equity-related tasks showed 10-20% error amplification with demographic modifications, and human reviewers identified clinically relevant failures missed by automated evaluation. Our findings demonstrate that performance variance and worst-case failures provide more clinically meaningful reliability indicators than mean accuracy alone, and that hybrid evaluation approaches combining automation with clinician oversight are essential for credible safety assessment.
David Talby, Andrei Marian Feier, Veysel Kocaman, Yigit Gul, Ahmet Korkmaz, Alexander Thomas, Aleksei Zakharov, Jay Gil · 9 authors totalSpecialty-Specific Medical Language Model for Immune-Mediated Diseases
arXiv · arXiv 2605.28838 · 0 citations · Source: openalexExtracting detailed clinical information from free-text medical narratives remains a practical challenge for researchers and healthcare systems. Terminology for immune-mediated and infectious diseases is especially inconsistent across sources, which often limits the ability of general-purpose Natural Language Processing (NLP) systems to capture the relevant biomedical concepts with sufficient granularity. We developed a domain-specific Named Entity Recognition (NER) model tailored to identify disease-related entities occurring in immunology and infectious disease contexts. We assembled and manually annotated a dataset of 371 case reports in collaboration with two clinical specialists, defining twelve entity classes covering immune-mediated and infectious conditions as well as related symptoms and clinical descriptors. We evaluated several modeling strategies, including the MedicalNER architecture with multiple healthcare-specific embeddings, a BERT-based token classification model, and zero-shot NER systems. The strongest performance was obtained with a transformer-based model trained on clinical-domain embeddings, which reached an F1 score of 0.89, consistently outperforming baseline and zero-shot approaches. The combination of specialized embeddings and expert annotation proved particularly valuable for capturing nuanced disease terminology and improving generalization across heterogeneous biomedical text. The prompted LLM baseline achieved substantially lower performance under the same evaluation protocol, reflecting difficulties in producing span-consistent outputs for fine-grained entity boundaries despite detailed prompting. The resulting model provides a structured way to analyze case reports and can support downstream tasks such as cohort identification, disease monitoring, and clinical decision support.
David Talby, Veysel Kocaman, Gürsev Pirge, Yigit Gul, Au Vo, Zhenya Nargizyan · 6 authors totalLACUNA: Safe Agents as Recursive Program Holes
arXiv 2605.28617 · 0 citations · Source: arxivLLM agents increasingly act by writing code, yet a split persists between the runtime that drives the agent and the code the model writes. The runtime owns the loop, context, and control flow, and the model has little say over any of them. Letting model-written code shape the runtime itself would make agents more expressive, but it would also sharpen safety problems. A model can be diverted by a prompt injection, call the wrong tool, or fail partway and leave an inconsistent state, and each such failure reaches further when the code shapes the runtime than when it expresses a single action. We present LACUNA, a programming model for agents that closes this split while preserving safety. Each agent action is a typed call $\texttt{agent[T](task)}$ that the LLM fills with code when execution reaches it, and the code is type-checked against the surrounding program before it runs. Because each action is accepted or rejected as a whole, a rejected one leaves the environment untouched, and its compiler diagnostics drive a retry. The same check also bounds which tools and data an action may use and how they flow. Our primitive expresses ReAct loops, sub-agents, skills, parallel decomposition, and multi-model planning as ordinary control flow. We evaluate LACUNA on a collection of test cases, BrowseComp-Plus, and $τ^2$-bench. On BrowseComp-Plus, $8.6\%$ of generations are rejected before execution, with 0.7 retries per query on average, and the agent reaches $27.1\%$ accuracy. On $τ^2$-bench, LACUNA solves $76.0\%$ of $392$ tasks across four domains with a capable model, on par with the baseline agent.
Martin Odersky, Yaoyu Zhao, Yichen Xu, Oliver Bračevac, Cao Nguyen Pham, Frank Zhengqing Wu · 6 authors totalBeyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows
arXiv.org · DOI 10.48550/arXiv.2605.24219 · arXiv 2605.24219 · 0 citations · Source: semantic-scholar+arxivLarge Language Models (LLMs) are increasingly deployed as autonomous agents that reason, use tools, and act over multiple steps. Yet most hallucination benchmarks still evaluate only the final output, missing failures that originate in intermediate Thought-Action-Observation steps. We present Trajel, a dataset and evaluation framework for auditing trajectory-level hallucinations in multi-agent industrial workflows. Trajel introduces a five-type hallucination taxonomy (factual, referential, logical, procedural, and scope-based) over expert-annotated agent traces from AssetOpsBench. We benchmark supervised detection models at the subtask, trajectory, and long-context levels. Our results show that the most common failure modes are missed by existing benchmarks, that nearly half of hallucinated trajectories involve multiple types at once, and that automated detectors with high binary accuracy still misclassify the subtlest types. Trajectory-aware detection significantly outperforms standard post-hoc verification, making taxonomy-grounded evaluation necessary for safer agentic deployment.
Santosh Borse, Harshada H. Badave, Andrea Gomez, H. Narahari, Sarah Carter, Vishwa Bhatt, Aishani Rachakonda, Shuxin Lin · 9 authors totalTubiFM: Unified Item, Carousel, and Search Ranking for Streaming Discovery
arXiv · arXiv 2605.23702 · 0 citations · Source: semantic-scholarPersonalized discovery systems often train separate models for item ranking, carousel ranking, and search, even though these tasks expose complementary signals from the same viewer journey: watches shape carousel and item ranking, search queries reveal intent even when they do not lead to a catalog match, and watch history helps interpret search as rewatching, continuation, or new discovery. We introduce the user story, a serialized representation that turns a user's cross-surface history - attributes, sessions, watch events with surface and carousel context, and search events - into a single token sequence. By interleaving pretrained language tokens with domain-specific event tokens, user stories let heterogeneous recommendation and search tasks be expressed as prompted next-token prediction over a shared grammar. TubiFM is one instantiation of this approach: a Llama 3.2 1B-based model trained on user stories and prompted to rank items, carousels, or search results without task-specific architectures. In offline evaluation, this single model outperforms specialist baselines across item, carousel, and search ranking. In online A/B tests, TubiFM significantly improves search total viewing time (TVT) by $+3.9\%$ and carousel TVT by $+0.30\%$. Item ranking is statistically neutral on TVT ($+0.14\%$), but matches a mature production stack; across all three tasks, TubiFM serves on L40S GPUs and reduces p99 ranking latency from 500ms to 200ms. These results show that shared user stories can improve discovery while simplifying ranking systems.
Mike Tamir, Alexandre Salle, Chenglei Niu, Suchismit Mahapatra, Xiaoxiao Chen, Suvash Sedhain, Yaqi Wang, Shervin Shahryari · 10 authors totalAutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery
arXiv.org · DOI 10.48550/arXiv.2605.23204 · arXiv 2605.23204 · 7 citations · Source: arxiv+semantic-scholarScientific research is being reshaped by AI systems that move beyond isolated assistance toward longer-horizon workflows spanning literature grounding, hypothesis generation, experimentation, validation, reporting, and revision. This shift marks a transition from task-level AI for science to workflow-level research automation. Yet current systems remain fragmented, differing in autonomy, domain scope, execution environment, validation mechanism, and human oversight, while still struggling with evidence preservation, reproducibility, weak-direction rejection, provenance tracking, cross-domain robustness, and accountable scientific closure. This survey examines these developments through AutoResearch, defined as the developmental spectrum of AI-powered scientific workflow automation. Within it, Vibe Research denotes the human-steered region of prompt-based assistance and human-verified execution, whereas emerging AI-led systems coordinate larger portions of the discovery loop without achieving robust autonomy. We analyze how research systems redistribute control, evidence, execution, validation, and accountability across workflows and organize the field around five workflow conditions: literature and research grounding; hypothesis formation and planning; experimentation and tool use; feedback, validation, and review; and reporting and knowledge communication. We further synthesize AI scientist systems, mixed-initiative co-research frameworks, benchmarks, domain deployments, and open-source infrastructures. Finally, we propose five evaluation dimensions--novelty, validity, impact, reliability, and provenance--and show that AutoResearch autonomy is domain-conditioned, being more credible in structured, executable, and rapidly verifiable settings but limited in embodied, delayed, heterogeneous, ethical, or institutionally accountable contexts.
Ran Xu, Guiyao Tie, Jiawen Shi, D. Song, Yixiao Huang, Ziji Sheng, Xueyang Zhou, Daizong Liu · 23 authors totalRuntime-Structured Task Decomposition for Agentic Coding Systems
arXiv · arXiv 2605.15425 · Source: arxiv+semantic-scholarAgentic coding systems increasingly use large language models (LLMs) for software engineering tasks such as debugging, root cause analysis, and code review. However, many existing systems encode task logic, execution flow, and output generation inside monolithic prompts. This design creates brittle behavior, limited debuggability, and high retry costs because failures often require rerunning the full workflow. We present runtime-structured task decomposition, an architectural approach in which task partitioning and execution flow are managed through executable control logic rather than prompt structure alone. LLMs are used only for focused judgment tasks, and outputs are validated against predefined schemas before downstream execution. We evaluate this approach on two software engineering workloads using three configurations: monolithic execution, static decomposition with fixed subtasks and no runtime branching, and runtime-structured decomposition. Each configuration was evaluated across 10 runs. Our results show that decomposition alone does not necessarily reduce retry cost. In the Kubernetes root cause analysis workload, the static decomposition baseline produced a retry cost of 1,632 +/- 145 tokens versus 904 +/- 17 tokens for the monolithic baseline because failures forced reruns of downstream subtasks. A similar pattern appeared in the multi-file debugging workload, where the static baseline consumed 933 tokens compared to 703 tokens for the monolithic system. The runtime-structured approach reran only failed subtasks, reducing retry costs to 436 +/- 132 tokens for root cause analysis and 460 tokens for debugging. Overall, the approach achieved up to 51.7% lower retry cost than monolithic systems and 73.2% lower retry cost than static decomposition baselines, improving efficiency, debuggability, and operational reliability in agentic coding systems.
Ruchi Mahindru, Shubhi Asthana, Bing Zhang, Chad DeLuca, Hima Patel · 5 authors totalFirst-Class Refinement Types for Scala
arXiv 2605.08369 · 1 citations · Source: arxivRefinement types -- types qualified with logical predicates -- have proven effective for lightweight verification in languages like Liquid Haskell, F*, and Dafny. However, in these systems refinements are either written in a separate specification language or treated as second-class annotations, disconnected from the host language's type system. This disconnect creates usability barriers: programmers must maintain two mental models, and refinements cannot interact with features like type inference, subtyping, or overloading. We present the design of first-class refinement types for Scala 3, where refinements are ordinary types that participate in subtyping, inference, and pattern matching alongside existing language features. We prove type soundness of a core calculus mechanized in Rocq, combining dependent function types, bounded polymorphism, positive equi-recursive types, union and intersection types, and refinement types under a partial-correctness semantics using a fuel-bounded definitional interpreter and semantic typing. Finally, we implement our design as a prototype extension of the Scala 3 compiler with a lightweight e-graph-based solver for predicate entailment.
Martin Odersky, Matt Bovel, Viktor Kunčak · 3 authors total