Papers.
Research connected to its authors, projects, companies, talks, events, and the rest of the graph.
Add a paper ↗Model-based design and automated validation of ARINC653 architectures
RSP · DOI 10.1109/RSP.2015.7416539 · 4 citations · Source: semantic-scholar+openalexSafety-Critical Systems as used in avionics systems are now extremely software-reliant. As these systems are life-or mission- critical, software must be carefully designed and certified according to stringent standards. One typical pitfalls of such project is the late detection of safety issues or bugs at integration time that impose to redo development steps. Model-Based Engineering aims at capturing system concerns with a specific notations and use models to drive the development process through all its phases - design, validation, implementation and ultimately, certification. Through a single consistent notation, such an approach would avoid undefined assumption and traditional hurdles due to informal, text-based, specifications. In this paper, we present recent contributions we pushed forward in the AADL architecture description language for the design and validation of Integrated Modular Avionics systems. First, we review modeling patterns to support abstractions for IMA systems. We then introduce capabilities to check all ARINC653 patterns are enforced at model-level. In addition, we review errror modeling and safety analysis capabilities towards the production of safety reports conforming to ARP4761 recommandations.
Julien Delange, Jérôme Hugues · 2 authors totalBig Data: Promises and Problems
Computer · DOI 10.1109/mc.2015.62 · 151 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Venkat N. Gudivada, Ricardo Baeza‐Yates, Vijay V. Raghavan · 4 authors totalBitmap-Based On-line Analytical Processing of Time Interval Data
DOI 10.1109/itng.2015.9 · 5 citations · Source: openalex+career-authorityPhilipp Meisen, Diane Keng, Tobias Meisen, Marco Recchioni, Sabina Jeschke · 5 authors totalMatrix regressor adaptive observers for battery management systems
International Symposium on Intelligent Control · DOI 10.1109/ISIC.2015.7307293 · 14 citations · Source: semantic-scholarAleksandar Kojic, Benjamin Jenkins, A. Annaswamy, A. Kojic · 4 authors totalModels for sustainability
International Smart Cities Conference · DOI 10.1109/ISC2.2015.7366221 · 20 citations · Source: crossref+semantic-scholarRosario Uceda-Sosa, Berta Cormenzana, M. Marinescu, Mónica Marrero, Sergio Mendoza, Salvador Rueda, Rosario A. Uceda-Sosa · 7 authors totalLow-current spin transfer torque MRAM
IEEE International Magnetics Conference · DOI 10.1109/INTMAG.2015.7157006 · 3 citations · Source: semantic-scholarAnthony Annunziata, D. Worledge, A. Annunziata, S. Brown, W. Chen, J. Harms, G. Hu, Y. Kim · 18 authors totalDeploying clouds in the Guifi community network
IM · DOI 10.1109/INM.2015.7140428 · Source: dblp+first-party-career-authorityJim Dowling, Roger Baig, Pau Escrich Garcia, Felix Freitag, Roc Meseguer, Agustí Moll, Leandro Navarro 0001, Ermanno Pietrosemoli · 11 authors totalDecentralized credit mining in P2P systems
DOI 10.1109/ifipnetworking.2015.7145334 · 1 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Mihai Capotă, Dick Epema · 3 authors totalSTT-MRAM with double magnetic tunnel junctions
International Electron Devices Meeting · DOI 10.1109/IEDM.2015.7409772 · 76 citations · Source: semantic-scholarAnthony Annunziata, G. Hu, J. Lee, J. Nowak, J. Sun, J. Harms, A. Annunziata, S. Brown · 20 authors totalGP-GPIS-OPT: Grasp planning with shape uncertainty using Gaussian process implicit surfaces and Sequential Convex Programming
ICRA 2015 · DOI 10.1109/icra.2015.7139882 · 59 citations · Source: openalexComputing grasps for an object is challenging when the object geometry is not known precisely. In this paper, we explore the use of Gaussian process implicit surfaces (GPISs) to represent shape uncertainty from RGBD point cloud observations of objects. We study the use of GPIS representations to select grasps on previously unknown objects, measuring grasp quality by the probability of force closure. Our main contribution is GP-GPIS-OPT, an algorithm for computing grasps for parallel-jaw grippers on 2D GPIS object representations. Specifically, our method optimizes an approximation to the probability of force closure subject to antipodal constraints on the parallel jaws using Sequential Convex Programming (SCP). We also introduce GPIS-Blur, a method for visualizing 2D GPIS models based on blending shape samples from a GPIS. We test the algorithm on a set of 8 planar objects with transparency, translucency, and specularity. Our experiments suggest that GP-GPIS-OPT computes grasps with higher probability of force closure than a planner that does not consider shape uncertainty on our test objects and may converge to a grasp plan up to 5.7×faster than using Monte-Carlo integration, a common method for grasp planning under shape uncertainty. Furthermore, initial experiments on the Willow Garage PR2 robot suggest that grasps selected with GP-GPIS-OPT are up to 90% more successful than those planned assuming a deterministic shape. Our dataset, code, and videos of our experiments are available at http://rll.berkeley.edu/icra2015grasping/.
Jur van den Berg, Jeffrey Mahler, Sachin Patil, Ben Kehoe, Matei Ciocarlie, Pieter Abbeel, Ken Goldberg · 7 authors totalToward asymptotically optimal motion planning for kinodynamic systems using a two-point boundary value problem solver
ICRA 2015 · DOI 10.1109/icra.2015.7139776 · 62 citations · Source: openalexWe present an approach for asymptotically optimal motion planning for kinodynamic systems with arbitrary nonlinear dynamics amid obstacles. Optimal sampling-based planners like RRT*, FMT*, and BIT* when applied to kinodynamic systems require solving a two-point boundary value problem (BVP) to perform exact connections between nodes in the tree. Two-point BVPs are non-trivial to solve, hence the prevalence of alternative approaches that focus on specific instances of kinodynamic systems, use approximate solutions to the two-point BVP, or use random propagation of controls. In this work, we explore the feasibility of exploiting recent advances in numerical optimal control and optimization to solve these two-point BVPs for arbitrary kinodynamic systems and how they can be integrated with existing optimal planning algorithms. We combine BIT* with a two-point BVP solver that uses sequential quadratic programming (SQP). We consider the problem of computing minimum-time trajectories. Since the duration of trajectories is not known a-priori, we include the time-step as part of the optimization to allow SQP to optimize over the duration of the trajectory while keeping the number of discrete steps fixed for every connection attempted. Our experiments indicate that using a two-point BVP solver in the inner-loop of BIT* is competitive with the state-of-the-art in sampling-based optimal planning that explicitly avoids the use of two-point BVP solvers.
Jur van den Berg, Christopher Xie, Sachin Patil, Pieter Abbeel · 4 authors totalThe First Facial Landmark Tracking in-the-Wild Challenge: Benchmark and Results
2015 IEEE International Conference on Computer Vision Workshop (ICCVW) · DOI 10.1109/ICCVW.2015.132 · 332 citations · Source: semantic-scholarDetection and tracking of faces in image sequences is among the most well studied problems in the intersection of statistical machine learning and computer vision. Often, tracking and detection methodologies use a rigid representation to describe the facial region 1, hence they can neither capture nor exploit the non-rigid facial deformations, which are crucial for countless of applications (e.g., facial expression analysis, facial motion capture, high-performance face recognition etc.). Usually, the non-rigid deformations are captured by locating and tracking the position of a set of fiducial facial landmarks (e.g., eyes, nose, mouth etc.). Recently, we witnessed a burst of research in automatic facial landmark localisation in static imagery. This is partly attributed to the availability of large amount of annotated data, many of which have been provided by the first facial landmark localisation challenge (also known as 300-W challenge). Even though now well established benchmarks exist for facial landmark localisation in static imagery, to the best of our knowledge, there is no established benchmark for assessing the performance of facial landmark tracking methodologies, containing an adequate number of annotated face videos. In conjunction with ICCV'2015 we run the first competition/challenge on facial landmark tracking in long-term videos. In this paper, we present the first benchmark for long-term facial landmark tracking, containing currently over 110 annotated videos,
Jean Kossaifi, Jie Shen, S. Zafeiriou, Grigorios G. Chrysos, Georgios Tzimiropoulos, M. Pantic · 6 authors totalA crosslinguistic study of prosodic focus
DOI 10.1109/icassp.2015.7178873 · 28 citations · Source: openalex+first-party-career-authorityMark Liberman, Yong-cheol Lee, Bei Wang, Sisi Chen, Martine Adda‐Decker, Angélique Amelot, Satoshi Nambu · 7 authors totalA security framework for population-scale genomics analysis
HPCS · DOI 10.1109/HPCSIM.2015.7237028 · Source: dblp+first-party-career-authorityJim Dowling, Ali Gholami, Erwin Laure · 3 authors totalAUDIME: Augmented disaster medicine
DOI 10.1109/healthcom.2015.7454522 · 4 citations · Source: openalex+career-authorityPhilipp Meisen, Alexander Paulus, Tobias Meisen, Sabina Jeschke, Michael Czaplik, F. Hirsch · 6 authors totalSimilarity Search of Bounded TIDASETs within Large Time Interval Databases
DOI 10.1109/csci.2015.36 · 4 citations · Source: openalex+career-authorityPhilipp Meisen, Diane Keng, Tobias Meisen, Marco Recchioni, Sabina Jeschke · 5 authors totalControl-oriented modeling and adaptive parameter estimation of a Lithium ion intercalation cell
IEEE Conference on Decision and Control · DOI 10.1109/CDC.2015.7403375 · 3 citations · Source: semantic-scholarAleksandar Kojic, Pierre Y. Bi, A. Annaswamy, A. Kojic · 4 authors totalKey-value store implementations for Arduino microcontrollers.
CCECE · DOI 10.1109/ccece.2015.7129178 · Source: dblp+ubc-authorityRamon Lawrence, Scott Fazackerley, Eric Huang, Graeme Douglas, Raffi Kudlac · 5 authors totalSentiment Expression via Emoticons on Social Media
2015 IEEE International Conference on Big Data (Big Data) · DOI 10.1109/BigData.2015.7364034 · arXiv 1511.02556 · 89 citations · Source: semantic-scholar+arxivEmoticons (e.g., :) and :( ) have been widely used in sentiment analysis and other NLP tasks as features to machine learning algorithms or as entries of sentiment lexicons. In this paper, we argue that while emoticons are strong and common signals of sentiment expression on social media the relationship between emoticons and sentiment polarity are not always clear. Thus, any algorithm that deals with sentiment polarity should take emoticons into account but extreme caution should be exercised in which emoticons to depend on. First, to demonstrate the prevalence of emoticons on social media, we analyzed the frequency of emoticons in a large recent Twitter data set. Then we carried out four analyses to examine the relationship between emoticons and sentiment polarity as well as the contexts in which emoticons are used. The first analysis surveyed a group of participants for their perceived sentiment polarity of the most frequent emoticons. The second analysis examined clustering of words and emoticons to better understand the meaning conveyed by the emoticons. The third analysis compared the sentiment polarity of microblog posts before and after emoticons were removed from the text. The last analysis tested the hypothesis that removing emoticons from text hurts sentiment classification by training two models with and without emoticons in the text, respectively. The results confirms the arguments that: 1) a few emoticons are strong and reliable signals of sentiment polarity and one should take advantage of them in any sentiment analysis; 2) a large group of the emoticons conveys complicated sentiment hence they should be treated with extreme caution.
Jorge Castanon, Hao Wang, Jorge A. Castanon · 3 authors totalKlout score: Measuring influence across multiple social networks
IEEE International Conference on Big Data (Big Data) · DOI 10.1109/BigData.2015.7364017 · arXiv 1510.08487 · 113 citations · Source: semantic-scholar+arxivIn this work, we present the Klout Score, an influence scoring system that assigns scores to 750 million users across 9 different social networks on a daily basis. We propose a hierarchical framework for generating an influence score for each user, by incorporating information for the user from multiple networks and communities. Over 3600 features that capture signals of influential interactions are aggregated across multiple dimensions for each user. The features are scalably generated by processing over 45 billion interactions from social networks every day, as well as by incorporating factors that indicate real world influence. Supervised models trained from labeled data determine the weights for features, and the final Klout Score is obtained by hierarchically combining communities and networks. We validate the correctness of the score by showing that users with higher scores are able to spread information more effectively in a network. Finally, we use several comparisons to other ranking systems to show that highly influential and recognizable users across different domains have high Klout scores.
Adithya Rao, Nemanja Spasojevic, Zhisheng Li, Trevor Dsouza · 4 authors totalIdentifying actionable messages on social media
IEEE International Conference on Big Data (Big Data) · DOI 10.1109/BigData.2015.7364016 · arXiv 1511.00722 · 9 citations · Source: semantic-scholar+arxivText actionability detection is the problem of classifying user authored natural language text, according to whether it can be acted upon by a responding agent. In this paper, we propose a supervised learning framework for domain-aware, large-scale actionability classification of social media messages. We derive lexicons, perform an in-depth analysis for over 25 text based features, and explore strategies to handle domains that have limited training data. We apply these methods to over 46 million messages spanning 75 companies and 35 languages, from both Facebook and Twitter. The models achieve an aggregate population-weighted F measure of 0.78 and accuracy of 0.74, with values of over 0.9 in some cases.
Adithya Rao, Nemanja Spasojevic · 2 authors totalScientific computing meets big data technology: An astronomy use case
IEEE International Conference on Big Data (Big Data) 2015 · DOI 10.1109/bigdata.2015.7363840 · arXiv 1507.03325 · 54 citations · Source: openalexScientific analyses commonly compose multiple single-process programs into a dataflow. An end-to-end dataflow of single-process programs is known as a many-task application. Typically, tools from the HPC software stack are used to parallelize these analyses. In this work, we investigate an alternate approach that uses Apache Spark - a modern big data platform - to parallelize many-task applications. We present Kira, a flexible and distributed astronomy image processing toolkit using Apache Spark. We then use the Kira toolkit to implement a Source Extractor application for astronomy images, called Kira SE. With Kira SE as the use case, we study the programming flexibility, dataflow richness, scheduling capacity and performance of Apache Spark running on the EC2 cloud. By exploiting data locality, Kira SE achieves a 3.7 χ speedup over an equivalent C program when analyzing a 1TB dataset using 512 cores on the Amazon EC2 cloud. Furthermore, we show that by leveraging software originally designed for big data infrastructure, Kira SE achieves competitive performance to the C implementation running on the NERSC Edison supercomputer. Our experience with Kira indicates that emerging Big Data platforms such as Apache Spark are a performant alternative for many-task scientific applications.
Evan R. Sparks, Frank Austin Nothaft, Zhao Zhang, Kyle Barbary, Evan Sparks, Oliver Zahn, Michael J. Franklin, David A. Patterson · 8 authors totalFuzzing the Rust Typechecker Using CLP (T)
ASE · DOI 10.1109/ASE.2015.65 · 82 citations · Source: semantic-scholar+dblpJared Roesch, Kyle Dewey, Ben Hardekopf · 3 authors totalHardware acceleration of Private Information Retrieval protocols using GPUs
IEEE International Conference on Application-Specific Systems, Architectures, and Processors · DOI 10.1109/ASAP.2015.7245719 · 5 citations · Source: semantic-scholarMihai Maruseac, Gabriel Ghinita, Ming Ouyang, R. Rughinis · 4 authors totalControl-oriented modeling of spark assisted compression ignition using a double Wiebe function
American Control Conference · DOI 10.1109/ACC.2015.7172076 · 3 citations · Source: semantic-scholarAleksandar Kojic, Z. Qu, N. Ravi, J. Oudart, E. Doran, V. Mittal, A. Kojic · 7 authors totalModel for Thermal Relic Dark Matter of Strongly Interacting Massive Particles
Phys.Rev.Lett. · DOI 10.1103/PhysRevLett.115.021301 · arXiv 1411.3727 · 455 citations · Source: inspirehep+author-first-partyJay Wacker, Yonit Hochberg, Eric Kuflik, Hitoshi Murayama, Tomer Volansky, Jay G. Wacker · 6 authors totalThe NIH BD2K center for big data in translational genomics
Journal of the American Medical Informatics Association · DOI 10.1093/jamia/ocv047 · 36 citations · Source: openalexFrank Austin Nothaft, Benedict Paten, Mark Diekhans, Brian Druker, Stephen Friend, Justin Guinney, N.C. Gassner, Mitchell Guttman · 20 authors totalHarmony Assumptions in Information Retrieval and Social Networks
The Computer Journal · DOI 10.1093/comjnl/bxv031 · 3 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Thomas Roelleke, Andreas Kaltenbrunner, Ricardo Baeza‐Yates · 4 authors totalValidation Study for Automated Nucleic Acid Extraction From Frozen Tissue, Peripheral Blood, and Formalin-Fixed Paraffin-Embedded Tissue
American Journal of Clinical Pathology · DOI 10.1093/ajcp/144.suppl2.240 · 0 citations · Source: openalex+authoritative-profilePetros Giannikopoulos, Jeff Catalano, Kimberly Lung, Mita Patel, Catherine K. Foo, Anibal Cordero · 6 authors totalPlacing a Face to a Case: Investigating the Impact of Nonconventional Laboratory Practices on Staff Morale and Job Satisfaction
American Journal of Clinical Pathology · DOI 10.1093/ajcp/144.suppl2.187 · 0 citations · Source: openalex+authoritative-profilePetros Giannikopoulos, Jeff Catalano, Catherine Foo, Mita Patel, Kimberly Lung, Anibal Cordero · 6 authors totalThe Automated Lab: Challenges, Triumphs, Pitfalls, and Considerations
American Journal of Clinical Pathology · DOI 10.1093/ajcp/144.suppl2.186 · 0 citations · Source: openalex+authoritative-profilePetros Giannikopoulos, Jeffrey G. Catalano, Mita Patel, Kimberly Lung, Catherine K. Foo, Anibal Cordero · 6 authors totalNoninvasive monitoring of infection and rejection after lung transplantation
Proceedings of the National Academy of Sciences · DOI 10.1073/pnas.1517494112 · 364 citations · Source: openalex+stanford-first-party+career-authorityLance Martin, Iwijn De Vlaminck, Michael A. Kertesz, K. Patel, Mark Kowarsky, Calvin Strehl, Garrett Cohen, Helen Luikart · 16 authors totalMaterials investigation for thermally-assisted magnetic random access memory robust against 400 °C temperatures
DOI 10.1063/1.4917066 · 9 citations · Source: semantic-scholarAnthony Annunziata, A. Annunziata, P. Trouilloud, S. Bandiera, Shelley Brown, E. Gapihan, E. O'Sullivan, D. Worledge · 8 authors totalIntrinsic retroviral reactivation in human preimplantation embryos and pluripotent cells
Nature · DOI 10.1038/nature14308 · 678 citations · Source: openalex+stanford-first-party+career-authorityLance Martin, Edward J. Grow, Ryan A. Flynn, Shawn L. Chavez, Nicholas Bayless, Mark Wossidlo, Daniel J. Wesche, Carol B. Ware · 12 authors totalThe psychology of intelligence analysis: Drivers of prediction accuracy in world politics.
Journal of Experimental Psychology Applied · DOI 10.1037/xap0000040 · 204 citations · Source: openalexThis article extends psychological methods and concepts into a domain that is as profoundly consequential as it is poorly understood: intelligence analysis. We report findings from a geopolitical forecasting tournament that assessed the accuracy of more than 150,000 forecasts of 743 participants on 199 events occurring over 2 years. Participants were above average in intelligence and political knowledge relative to the general population. Individual differences in performance emerged, and forecasting skills were surprisingly consistent over time. Key predictors were (a) dispositional variables of cognitive ability, political knowledge, and open-mindedness; (b) situational variables of training in probabilistic reasoning and participation in collaborative teams that shared information and discussed rationales (Mellers, Ungar, et al., 2014); and (c) behavioral variables of deliberation time and frequency of belief updating. We developed a profile of the best forecasters; they were better at inductive reasoning, pattern detection, cognitive flexibility, and open-mindedness. They had greater understanding of geopolitics, training in probabilistic reasoning, and opportunities to succeed in cognitively enriched team environments. Last but not least, they viewed forecasting as a skill that required deliberate practice, sustained effort, and constant monitoring of current affairs.
Lyle Ungar, Barbara A. Mellers, Eric Stone, Pavel Atanasov, Nick Rohrbaugh, S. Emlen Metz, Michaël Bishop, Michael C. Horowitz · 10 authors totalA representation theorem for second-order functionals
Journal of Functional Programming · DOI 10.1017/s0956796815000088 · 1 citations · Source: openalex+career-authorityRussell O'Connor, Mauro Jaskelioff, Russell O’Connor · 3 authors totalCUFP'13 scribe's report
Journal of functional programming · DOI 10.1017/S0956796815000052 · 0 citations · Source: semantic-scholarMarius Eriksen, Michael Sperber, Anil Madhavapeddy · 3 authors totalFlow simulation and analysis of high-power flow batteries
DOI 10.1016/J.JPOWSOUR.2015.08.041 · 75 citations · Source: semantic-scholarAleksandar Kojic, E. Knudsen, P. Albertus, K. Cho, A. Weber, A. Kojic · 6 authors totalTime and information retrieval: Introduction to the special issue
Inf. Process. Manag. · DOI 10.1016/j.ipm.2015.05.002 · 16 citations · Source: dblp+semantic-scholarOmar Alonso, Leon Derczynski, Jannik Strötgen, Ricardo Campos · 4 authors totalPlaying with knowledge: A virtual player for “Who Wants to Be a Millionaire?” that leverages question answering techniques
Artificial Intelligence · DOI 10.1016/j.artint.2015.02.003 · 20 citations · Source: openalexPiero Molino, Pasquale Lops, Giovanni Semeraro, Marco de Gemmis, Pierpaolo Basile · 5 authors totalEffects of a Community-Based Fall Management Program on Medicare Cost Savings
American Journal of Preventive Medicine · DOI 10.1016/j.amepre.2015.07.004 · Source: publisher+pubmed+career-authorityDaniella Perlroth, E. Ghimire, E. M. Colligan, B. Howell, G. Marrufo, E. Rusev, M. Packard · 7 authors totalMeasuring Inter-site Engagement in a Network of Sites
Handbook of statistics · DOI 10.1016/b978-0-444-63492-4.00013-7 · 0 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Janette Lehmann, Mounia Lalmas, Ricardo Baeza‐Yates · 4 authors totalContinuous Model Selection for Large-Scale Recommender Systems
DOI 10.1016/B978-0-444-63492-4.00005-8 · 24 citations · Source: semantic-scholarSimon Chan, P. Treleaven · 2 authors totalHow to use NodeXL
Elsevier eBooks · DOI 10.1016/b978-0-12-801656-5.00022-6 · 0 citations · Source: openalex+first-party-career-authorityMarc Smith, Derek Hansen, Marc A. Smith · 3 authors totalRobust semantic text similarity using LSA, machine learning, and linguistic resources
Language Resources and Evaluation · DOI 10.1007/s10579-015-9319-2 · 41 citations · Source: semantic-scholarAbhay Kashyap, Abhay L. Kashyap, Lushan Han, Roberto Yus, Jennifer Sleeman, Taneeya Satyapanich, Sunil Gandhi, Tim Finin · 8 authors totalHow to present more readable text for people with dyslexia
Universal Access in the Information Society · DOI 10.1007/s10209-015-0438-8 · 104 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Luz Rello, Ricardo Baeza‐Yates · 3 authors totalPaying the Guard: An Entry-Guard-Based Payment System for Tor
Lecture notes in computer science · DOI 10.1007/978-3-662-47854-7_26 · 6 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Paolo Palmieri · 2 authors totalFirenzina: Porting a Chess Engine to Android
Lecture notes in electrical engineering · DOI 10.1007/978-3-662-47669-7_19 · 2 citations · Source: openalexDmitri Gusev, Corey Abshire, Dmitri A. Gusev · 3 authors totalBiobankCloud: A Platform for the Secure Storage, Sharing, and Processing of Large Biomedical Data Sets
Big-O(Q)/DMAH@VLDB · DOI 10.1007/978-3-319-41576-5_7 · Source: dblp+first-party-career-authorityJim Dowling, Alysson Bessani, Jörgen Brandt, Marc Bux, Vinicius Vielmo Cogo, Lora Dimitrova, Ali Gholami, Kamal Hakimzadeh · 17 authors totalReviewer Integration and Performance Measurement for Malware Detection
International Conference on Detection of intrusions and malware, and vulnerability assessment · DOI 10.1007/978-3-319-40667-1_7 · arXiv 1510.07338 · 89 citations · Source: semantic-scholarWe present and evaluate a large-scale malware detection system integrating machine learning with expert reviewers, treating reviewers as a limited labeling resource. We demonstrate that even in small numbers, reviewers can vastly improve the system's ability to keep pace with evolving threats. We conduct our evaluation on a sample of VirusTotal submissions spanning 2.5i?źyears and containing 1.1 million binaries with 778i?źGB of raw feature data. Without reviewer assistance, we achieve 72i?ź% detection at a 0.5i?ź% false positive rate, performing comparable to the best vendors on VirusTotal. Given a budget of 80 accurate reviews daily, we improve detection to 89i?ź% and are able to detect 42i?ź% of malicious binaries undetected upon initial submission to VirusTotal. Additionally, we identify a previously unnoticed temporal inconsistency in the labeling of training datasets. We compare the impact of training labels obtained at the same time training data is first seen with training labels obtained months later. We find that using training labels obtained well after samples appear, and thus unavailable in practice for current training data, inflates measured detection by almost 20i?ź% points. We release our cluster-based implementation, as well as a list of all hashes in our evaluation and 3i?ź% of our entire dataset.
Vaishaal Shankar, Brad Miller, Alex Kantchelian, Michael Carl Tschantz, Sadia Afroz, Rekha Bachwani, Riyaz Faizullabhoy, Ling Huang · 12 authors total