Papers.
Research connected to its authors, projects, companies, talks, events, and the rest of the graph.
Add a paper ↗Does Personalization Benefit Everyone in the Same Way? Multilingual Search Personalization for English vs. Non-English Users
UMAP Workshops · Source: dblp+adapt-autodesk-authorityAlex O'Connor, M. Rami Ghorab, Séamus Lawless, Alexander O'Connor, Vincent Wade · 5 authors totalDistributed Representations for Semantic Matching in non-factoid Question Answering
6 citations · Source: openalexUsers ’ interactions with search engines is shifting towards more complex information needs and the need for a deeper semantic un-derstanding of the query intent is needed. In this paper we propose a novel semantic matching criterion that adopts distributed representations of words in order to address com-plex information needs in a scalable way. We show that combining this criterion with other well established features it is possible to obtain over 22 % improvement for MRR and 27 % in P@1 over the best performing approach for answer-ing non-factoid questions, a specific form of complex information need. Moreover we show that in our setting our criterion can sub-stitute more complex linguistic feature.
Piero Molino, Luca Maria Aiello · 2 authors totalCoordination-Avoiding Database Systems.
arXiv (Cornell University) · 13 citations · Source: openalex+authoritative-profilePeter Bailis, Alan Fekete, Michael J. Franklin, Ali Ghodsi, Joseph M. Hellerstein, Ion Stoica · 6 authors totalConvolutional nets and watershed cuts for real-time semantic labeling of RGBD videos
Journal of Machine Learning Research (JMLR), vol. 15 · Source: dblpThis work addresses multi-class segmentation of indoor scenes with RGB-D inputs. While this area of research has gained much attention recently, most works still rely on handcrafted features. In contrast, we apply a multiscale convolutional network to learn features directly from the images and the depth information. Using a frame by frame labeling, we obtain nearly state-of-the-art performance on the NYU-v2 depth data set with an accuracy of 64.5%. We then show that the labeling can be further improved by exploiting the temporal consistency in the video sequence of the scene. To that goal, we present a method producing temporally consistent superpixels from a streaming video. Among the different methods producing superpixel segmentations of an image, the graph-based approach of Felzenszwalb and Huttenlocher is broadly employed. One of its interesting properties is that the regions are computed in a greedy manner in quasi-linear time by using a minimum spanning tree. In a framework exploiting minimum spanning trees all along, we propose an efficient video segmentation approach that computes temporally consistent pixels in a causal manner, filling the need for causal and real-time applications. We illustrate the labeling of indoor scenes in video sequences that could be processed in real-time using appropriate hardware such as an FPGA.
Clément Farabet, Camille Couprie, Laurent Najman, Yann LeCun · 4 authors totalCase-Based Subgoaling in Real-Time Heuristic Search for Video Game Pathfinding.
CoRR · Source: dblp+ubc-authorityRamon Lawrence, Vadim Bulitko, Yngvi Björnsson · 3 authors totalAutomatic Optimization for Data-Parallel Streaming Systems
German Research Training Groups in Computer Science Workshop · Source: dblpMatthias Sax, Matthias J. Sax · 2 authors totalAdditional Material for "Unifying Data Representation Transformations"
3 citations · Source: semantic-scholarVlad Ureche · 1 author totalIs Security Realistic in Cloud Computing?
Journal of International Technology and Information Management · DOI 10.58729/1941-6679.1020 · 26 citations · Source: openalex+semantic-scholarINTRODUCTION Cloud computing today is benefiting from the technological advancements in communication, storage and computing. The basic idea in cloud computing is to take advantage of economies of scale if IT services could be provided on demand with a decentralized infrastructure. This idea is a natural evolution from the IT time-share model of the 1960s and 1970s. Today, technology has advanced significantly and many more organizations have computing demands that are elastic in nature. Organizations large and small require reliable computing resources in order to succeed in business. Large businesses deal with complex systems where as Small and Medium sized Enterprises (SMEs) need access to affordable computing resources. Based on these aspects we can summarize some of the rationale for today's cloud computing needs as follows: * acquiring and managing the IT resources requires specialized skills, * maintaining a reliable IT infrastructure is expensive, * rapid technology advancements make it difficult to keep current the IT expertise, * internet has opened up many opportunities for individuals as well as small businesses, * number of entities requiring computing resources has grown exponentially, * SMEs' demand for computing resources varies significantly over time, * providing data security is a complex undertaking. In the above paragraph we have identified some of the major reasons as to why cloud computing would be advantageous to use. When a significant part of the business depends on a type of service that the business does not fully control, the question arises as to how the business can meet its obligations to its customers. As highlighted above, IT services are essential to the success of the business but it would be cost prohibitive for the business to manage an IT center with the required expertise and fluctuating demand on resources for processing and storage. Thus, a business using cloud computing must understand the security challenges that it would
S. Srinivasan, Jesse H. Jones · 2 authors totalAre Multi-way Joins Actually Useful?.
ICEIS (1) · DOI 10.5220/0004412100130022 · Source: dblp+ubc-authorityRamon Lawrence, Michael Henderson 0001 · 2 authors totalThe Simulated Greedy Algorithm for Several Submodular Matroid Secretary Problems.
STACS · DOI 10.4230/LIPIcs.STACS.2013.478 · Source: dblp+stanford-authorityTengyu Ma, Tengyu Ma 0001, Bo Tang 0003, Yajun Wang 0001 · 4 authors totalIT Landmarks in Less-Developed Countries: The Chilean Case
Proceedings of the Annual Conference of CAIS / Actes du congrès annuel de l ACSI · DOI 10.29173/cais140 · 2 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates, David Fuller, José A. Pino · 4 authors totalDigital Forensics Curriculum in Security Education
Journal of Information Technology Education Innovations in Practice · DOI 10.28945/1857 · 14 citations · Source: openalex+semantic-scholarS. Srinivasan, Linda V. Knight · 2 authors totalThe Relation between CEO Compensation and Past Performance
DOI 10.2308/ACCR-50274 · 144 citations · Source: semantic-scholarJose Plehn, R. Banker, M. Darrough, Ronghong Huang, J. Plehn-Dujowich · 5 authors totalSnowmass Energy Frontier Simulations
Peer-reviewed physics publication · DOI 10.2172/1128171 · arXiv 1309.1057 · 101 citations · Source: inspirehep+author-first-partyJay Wacker, Aram Avetisyan, Jacob Anderson, Jay G. Wacker, and collaborators · 5 authors totalMethods and Results for Standard Model Event Generation at $\sqrt{s}$ = 14 TeV, 33 TeV and 100 TeV Proton Colliders (A Snowmass Whitepaper)
Peer-reviewed physics publication · DOI 10.2172/1128125 · arXiv 1308.1636 · 108 citations · Source: inspirehep+author-first-partyJay Wacker, Aram Avetisyan, John M. Campbell, Timothy Cohen, Nitish Dhingra, James Hirschauer, Kiel Howe, Sudhir Malik · 12 authors totalIncremental learning for automated knowledge capture.
Sandia National Laboratories (SNL) · DOI 10.2172/1121921 · 0 citations · Source: semantic-scholar+openalexPeople responding to high-consequence national-security situations need tools to help them make the right decision quickly. The dynamic, time-critical, and ever-changing nature of these situations, especially those involving an adversary, require models of decision support that can dynamically react as a situation unfolds and changes. Automated knowledge capture is a key part of creating individualized models of decision making in many situations because it has been demonstrated as a very robust way to populate computational models of cognition. However, existing automated knowledge capture techniques only populate a knowledge model with data prior to its use, after which the knowledge model is static and unchanging. In contrast, humans, including our national-security adversaries, continually learn, adapt, and create new knowledge as they make decisions and witness their effect. This artificial dichotomy between creation and use exists because the majority of automated knowledge capture techniques are based on traditional batch machine-learning and statistical algorithms. These algorithms are primarily designed to optimize the accuracy of their predictions and only secondarily, if at all, concerned with issues such as speed, memory use, or ability to be incrementally updated. Thus, when new data arrives, batch algorithms used for automated knowledge capture currently require significant recomputation, frequently from scratch, which makes them ill suited for use in dynamic, timecritical, high-consequence decision making environments. In this work we seek to explore and expand upon the capabilities of dynamic, incremental models that can adapt to an ever-changing feature space.
Justin Basilico, Zachary Benz, Warren Davis, Kevin R. Dixon, Brian K. Jones, Nathaniel G. Martin, Jeremy Wendt · 7 authors totalAutomatic phonetic segmentation using boundary models
DOI 10.21437/interspeech.2013-540 · 60 citations · Source: openalex+first-party-career-authorityMark Liberman, Jiahong Yuan, Neville Ryant, Andreas Stolcke, Vikramjit Mitra, Wen Wang · 6 authors totalSpeech activity detection on youtube using deep neural networks
DOI 10.21437/interspeech.2013-203 · 102 citations · Source: openalex+first-party-career-authorityMark Liberman, Neville Ryant, Jiahong Yuan · 3 authors totalChina's Impending Patent Valuation Crisis: How the CCP's IP Talisman Dilutes Patent Value
SSRN working paper · DOI 10.2139/ssrn.2359913 · Source: semantic-scholar+crossrefNicole Shanahan, Nicole Ann Shanahan · 2 authors totalDeconstructing the Patent Bubble: An Exploration of Patent Monetization Entities from Sewing Machine Combination to Rockstar Bid Co.
SSRN working paper · DOI 10.2139/ssrn.2359912 · Source: semantic-scholar+crossrefThe question I address in this paper is whether or not the patent marketplace is developing in a sustainable way that rewards inventors and business people according to the aims of past and current patent policy. This paper addresses developments in the private patent marketplace from a historical perspective, and focuses on recent innovations in monetization models. The companies covered in this paper have all risen from the relative freedom under which the government has allowed them to operate, and their stories as well as their struggles explain the trends we see today and provide insight into where the patent marketplace is headed tomorrow.
Nicole Shanahan · 1 author totalA Comparison of the United States and European Fee-Shifting Standards and the Impact of the Proposed Trans-Pacific Partnership FTA on National Fee Shifting Law
SSRN working paper · DOI 10.2139/ssrn.2359910 · Source: semantic-scholar+crossrefFee shifting laws are designed to decrease the number of weak, nuisance lawsuits. However, patent-specific fee-shifting laws dramatically differ between countries and even jurisdictions within the U.S. This paper first analyzes the statutory foundations of patent fee shifting laws in the U.S. against select European jurisdictions that practice the "loser pays all" model. It then provides an empirical comparison of the decision signals such laws send to patent holders, and the affect these signals have in terms of the frequency of nuisance suits, especially those brought by non-practicing entities (NPEs). Finally, this paper evaluates how the proposed TPP, in its currently available version, would affect the proposed compulsory fee-shifting SHIELD Act (Saving High-tech Innovators from Egregious Legal Disputes) in the U.S. and the equivalent free shifting laws in Europe.
Nicole Shanahan · 1 author totalThe Policy Issue Surrounding Smart-Phone SEPs, and a Comment on Adjusting Current Patent Laws to Facilitate Clean-Tech Industry Growth
SSRN working paper · DOI 10.2139/ssrn.2359908 · Source: semantic-scholar+crossrefWhile the FTC has weighed in on the SEP debate, a closer look at the smart-phone battle unveils a situation that is not so clear-cut. This paper addresses the operations of SSO’s and the current issue with injunctive relief, and then explains why the Smartphone War presents unique questions of law. The paper then closes with a comment on significant smart-phone cases both completed and pending.
Nicole Shanahan · 1 author totalClearing the Air for Efficient Spectrum: A Hybrid Approach Towards Market Driven Reallocation in 2014
SSRN working paper · DOI 10.2139/ssrn.2359906 · Source: semantic-scholar+crossrefThis paper addresses the impact that current FCC spectrum policy has on spectrum distribution, and why current trends in media usage mandate changes to this system that do not include the government’s current repacking plans. The paper argues that the FCC should remove the current FCC allocation system based on communication categories, and implement in its place, a regulatory system that addresses high-level networks with aggregated distribution platforms, and utilize an "open wireless" model where possible to promote innovation and band efficiency. The boundaries on exclusive licenses should be structured around consumption metrics, i.e. bits exchanged, resulting as the sum of all parts of the transferred data. The current controversy over the repacking of the TV bands supports the argument that the FCC should make these proposed changes sooner rather than later to avoid disruption. Several recent laws have given the FCC greater authority to reallocate some of the TV bands to support new wireless data services. This process of repackaging is proving difficult given certain inflexibilities in the current spectrum classification scheme that allocates spectrum based on media platform. As Internet communications have changed the methods by which consumers use media, for example, a TV can be a means of accessing radio, the computer can be a TV source, and the Internet serves most telephone needs, the ability to correctly classify spectrum using traditional means is ineffective. This has resulted in the inefficient distribution of spectrum, and encumbered "white spaces" during a time when spectrum available for wireless broadband applications is under increasing strain.
Nicole Shanahan · 1 author totalSanta Clara Law Best Practices in Patent Litigation Survey
SSRN working paper · DOI 10.2139/ssrn.2321363 · Source: semantic-scholar+crossrefOver the past few years, Congress, appellate, and district courts have made significant strides to improve patent law and litigation practice. Congress is now considering making more changes, to supplement ongoing tailoring by the courts. Dialog between the patent bench, patent bar, and lawmakers is crucial for informing these efforts. To support this dialog, we developed, in consultation with judges and company lawyers in spring of 2013, a list of questions to probe the experiences, opinions, and suggestions of lawyers. We asked survey takers to rate, on a range from ineffective to very effective at enhancing the efficiency of litigation, certain existing and proposed practices and interventions, and converted these scores into numerical ratings (up to a highest possible effectiveness score of 100%). Based on over 500 responses, about a quarter from in-house counsel mostly at large technology companies and the remainder from outside (law firm) counsel, we probed a number of topics, and made a number of findings. For example, the highest rated intervention of any was timely decisions on summary judgment motions (86%) followed by timely decision on transfer motions (71%). Early claim construction also rated highly (around 68%). Among recent reforms, the FCAC e-Discovery model order ranked the highest in effectiveness (45%). However commentators said of many recent reforms that too much variance in court uptake and implementations, due to the discretion given to judges, undermined their effectiveness. Among proposed legislative reforms, fee-shifting and sanctions for prevailing parties and for discovery abuses rated most favorably (~65%). Based on about one hundred outside counsel responses, discovery abuses, followed by frivolous claims/defenses, were in the greatest need of sanction or shifting. Respondents also identified abuses that tended to be particular to plaintiffs: evasive discovery responses, overly burdensome or excessive discovery requests, frivolous/m
Nicole Shanahan, Colleen V. Chien, Daniel Dobkin, Wesley Helmholz, Coryn Millslagle, John Neal, Christopher Tosetti · 7 authors totalSupporting Social Data Observatory with Customizable Index Structures on HBase - Architecture and Performance
DOI 10.21236/ada603195 · 0 citations · Source: openalex+publisher+career-authorityKarissa McKelvey, Xiaoming Gao, Judy Qiu, Evan Roth, Clayton A. Davis, Andrew Younge, Emilio Ferrara, Fil Menczer · 8 authors totalA Multi-Teraflop Constituency Parser using GPUs
Conference on Empirical Methods in Natural Language Processing · DOI 10.18653/v1/d13-1195 · 22 citations · Source: semantic-scholarConstituency parsing with rich grammars remains a computational challenge. Graphics Processing Units (GPUs) have previously been used to accelerate CKY chart evaluation, but gains over CPU parsers were modest. In this paper, we describe a collection of new techniques that enable chart evaluation at close to the GPU’s practical maximum speed (a Teraflop), or around a half-trillion rule evaluations per second. Net parser performance on a 4-GPU system is over 1 thousand length30 sentences/second (1 trillion rules/sec), and 400 general sentences/second for the Berkeley Parser Grammar. The techniques we introduce include grammar compilation, recursive symbol blocking, and cache-sharing.
David Hall, J. Canny, David Leo Wright Hall, D. Klein · 4 authors totalRecursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
Conference on Empirical Methods in Natural Language Processing · DOI 10.18653/v1/d13-1170 · 9,634 citations · Source: semantic-scholarSemantic word spaces have been very useful but cannot express the meaning of longer phrases in a principled way. Further progress towards understanding compositionality in tasks such as sentiment detection requires richer supervised training and evaluation resources and more powerful models of composition. To remedy this, we introduce a Sentiment Treebank. It includes fine grained sentiment labels for 215,154 phrases in the parse trees of 11,855 sentences and presents new challenges for sentiment compositionality. To address them, we introduce the Recursive Neural Tensor Network. When trained on the new treebank, this model outperforms all previous methods on several metrics. It pushes the state of the art in single sentence positive/negative classification from 80% up to 85.4%. The accuracy of predicting fine-grained sentiment labels for all phrases reaches 80.7%, an improvement of 9.7% over bag of features baselines. Lastly, it is the only model that can accurately capture the effects of negation and its scope at various tree levels for both positive and negative phrases.
Richard Socher, R. Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, A. Ng, Christopher Potts · 8 authors totalRule-Based Information Extraction is Dead! Long Live Rule-Based Information Extraction Systems!
EMNLP · DOI 10.18653/v1/d13-1079 · 308 citations · Source: semantic-scholar+dblpThe rise of “Big Data” analytics over unstructured text has led to renewed interest in information extraction (IE). We surveyed the landscape of IE technologies and identified a major disconnect between industry and academia: while rule-based IE dominates the commercial world, it is widely regarded as dead-end technology by the academia. We believe the disconnect stems from the way in which the two communities measure the benefits and costs of IE, as well as academia’s perception that rulebased IE is devoid of research challenges. We make a case for the importance of rule-based IE to industry practitioners. We then lay out a research agenda in advancing the state-of-theart in rule-based IE systems which we believe has the potential to bridge the gap between academic research and industry practice.
Frederick Reiss, Laura Chiticariu, Yunyao Li · 3 authors totalFaster Optimal Planning with Partial-Order Pruning
International Conference on Automated Planning and Scheduling · DOI 10.1609/icaps.v23i1.13562 · 13 citations · Source: semantic-scholarWhen planning problems have many kinds of resources or high concurrency, each optimal state has exponentially many minor variants, some of which are "better" than others. Standard methods like \Astar cannot effectively exploit these minor relative differences, and therefore must explore many redundant, clearly suboptimal plans. We describe a new optimal search algorithm for planning that leverages a partial order relation between states. Under suitable conditions, states that are dominated by other states with respect to this order can be pruned while provably maintaining optimality. We also describe a simple method for automatically discovering compatible partial orders in both serial and concurrent domains. In our experiments we find that more than 98% of search states can be pruned in some domains.
David Hall, David Leo Wright Hall, Alon Cohen, David Burkett, D. Klein · 5 authors totalTrading Space for Time in Grid-Based Path Finding.
AAAI · DOI 10.1609/aaai.v27i1.8528 · Source: dblp+ubc-authorityRamon Lawrence, William Lee · 2 authors totalToward Interactive Grounded Language Acqusition
Robotics - Science and Systems · DOI 10.15607/RSS.2013.IX.005 · Source: dblp+author-first-party+semantic-machines-career-authorityJayant Krishnamurthy, Thomas Kollar, Grant P. Strimel · 3 authors totalHighly Available Transactions: Virtues and Limitations
PVLDB (VLDB 2014) · DOI 10.14778/2732232.2732237 · arXiv 1302.0309 · 257 citations · Source: semantic-scholarTo minimize network latency and remain online during server failures and network partitions, many modern distributed data storage systems eschew transactional functionality, which provides strong semantic guarantees for groups of multiple operations over multiple data items. In this work, we consider the problem of providing Highly Available Transactions (HATs): transactional guarantees that do not suffer unavailability during system partitions or incur high network latency. We introduce a taxonomy of highly available systems and analyze existing ACID isolation and distributed data consistency guarantees to identify which can and cannot be achieved in HAT systems. This unifies the literature on weak transactional isolation, replica consistency, and highly available systems. We analytically and experimentally quantify the availability and performance benefits of HATs--often two to three orders of magnitude over wide-area networks--and discuss their necessary semantic compromises.
Aaron Davidson, Peter Bailis, Alan Fekete, Ali Ghodsi, Joseph M. Hellerstein, Ion Stoica · 6 authors totalOn the Embeddability of Random Walk Distances
Proceedings of the VLDB Endowment · DOI 10.14778/2556549.2556554 · 36 citations · Source: dblp+semantic-scholarAnalysis of large graphs is critical to the ongoing growth of search engines and social networks. One class of queries centers around node affinity, often quantified by random-walk distances between node pairs. This paper studies whether random-walk distances can be embedded into a coordinate space for constant-time queries.
Adelbert Chang, Xiaohan Zhao, Atish Das Sarma, Haitao Zheng, Ben Y. Zhao · 5 authors totalSemantic Models for Re-ranking in Question Answering
Electronic workshops in computing · DOI 10.14236/ewic/fdia2013.5 · 1 citations · Source: openalexThis paper describes a research aimed at unveiling the role of Semantic Models in Question Answering. In these systems questions and answers are often expressed in quite different languages, so our objective is to bridge this “lexical chasm” adopting semantic representations. The aim of the research is to find out if Semantc Models are useful for this task and if they can improve the answer re-ranking performance. We have carried out an initial evaluation of a subset of the semantic models on the CLEF2010 QA dataset, showing their effectiveness. We also did a first attempt in combining them by means of Learning to Rank algorithms.
Piero Molino · 1 author totalMore Tweets, More Votes: Social Media as a Quantitative Indicator of Political Behavior
PLoS ONE · DOI 10.1371/journal.pone.0079449 · 299 citations · Source: openalex+publisher+career-authorityKarissa McKelvey, Joseph DiGrazia, Johan Bollen, Fabio Rojas · 4 authors totalChange in BMI Accurately Predicted by Social Exposure to Acquaintances
PLoS ONE · DOI 10.1371/journal.pone.0079238 · 10 citations · Source: semantic-scholar+dblp+career-authoritySai Moturu, Rahman O. Oloritun, T. Ouarda, S. Moturu, Anmol Madan, A. Pentland, Inas S. Khayal · 7 authors totalPersonality, Gender, and Age in the Language of Social Media: The Open-Vocabulary Approach
PLoS ONE · DOI 10.1371/journal.pone.0073791 · 1,747 citations · Source: openalexWe analyzed 700 million words, phrases, and topic instances collected from the Facebook messages of 75,000 volunteers, who also took standard personality tests, and found striking variations in language with personality, gender, and age. In our open-vocabulary technique, the data itself drives a comprehensive exploration of language that distinguishes people, finding connections that are not captured with traditional closed-vocabulary word-category analyses. Our analyses shed new light on psychosocial processes yielding results that are face valid (e.g., subjects living in high elevations talk about the mountains), tie in with other research (e.g., neurotic people disproportionately use the phrase 'sick of' and the word 'depressed'), suggest new hypotheses (e.g., an active life implies emotional stability), and give detailed insights (males use the possessive 'my' when mentioning their 'wife' or 'girlfriend' more often than females use 'my' with 'husband' or 'boyfriend'). To date, this represents the largest study, by an order of magnitude, of language and personality.
Lyle Ungar, H. Andrew Schwartz, Johannes C. Eichstaedt, Margaret L. Kern, Lukasz Dziurzynski, Stephanie M. Ramones, Megha Agrawal, Achal Shah · 11 authors totalThe Geospatial Characteristics of a Social Movement Communication Network
PLoS ONE · DOI 10.1371/journal.pone.0055957 · 156 citations · Source: openalex+publisher+career-authorityKarissa McKelvey, Michael Conover, Clayton A. Davis, Emilio Ferrara, Filippo Menczer, Alessandro Flammini · 6 authors totalGenomic Evidence for Island Population Conversion Resolves Conflicting Theories of Polar Bear Evolution
PLOS Genetics · DOI 10.1371/JOURNAL.PGEN.1003345 · Source: orcidJohn St. John, Cahill, James A., Green, Richard E., Fulton, Tara L., Stiller, Mathias, Jay, Flora, Ovsyanikov, Nikita, Salamzade, Rauf · 11 authors totalIntegrated genomic analysis of EGFR-mutant non-small cell lung cancer immediately following erlotinib initiation in patients.
Journal of Clinical Oncology · DOI 10.1200/jco.2013.31.15_suppl.11067 · 1 citations · Source: openalex+authoritative-profilePetros Giannikopoulos, Trever G. Bivona, Carlota Costa, Niki Karachaliou, Santiago Viteri, M.R. García-Campelo, John St. John, Andrew Uzilov · 19 authors totalROR1 mRNA expression in EGFR-mutant non-small-cell lung cancer (NSCLC) patients (p) with the T790M mutation: A potential therapeutic target.
Journal of Clinical Oncology · DOI 10.1200/jco.2013.31.15_suppl.11027 · 0 citations · Source: openalex+authoritative-profilePetros Giannikopoulos, Niki Karachaliou, Ana Drozdowskyj, Carlota Costa, Miguel Ángel Molina‐Vila, Ana Giménez‐Capitán, A. Vergnenègre, Bartomeu Massutí · 21 authors totalIntegrated genomic analysis by whole exome and transcriptome sequencing of tumor samples from EGFR-mutant non-small-cell lung cancer (NSCLC) patients (p) with acquired resistance to erlotinib.
Journal of Clinical Oncology · DOI 10.1200/jco.2013.31.15_suppl.11010 · 6 citations · Source: openalex+authoritative-profilePetros Giannikopoulos, Jonathan S. Weissman, John St. John, Andrew Uzilov, Carlota Costa, Niki Karachaliou, Irene Sansano, Eloísa Jantus‐Lewintre · 21 authors totalThe western painted turtle genome, a model for the evolution of extreme physiological adaptations in a slowly evolving lineage
Genome Biology · DOI 10.1186/GB-2013-14-3-R28 · Source: orcidJohn St. John, Shaffer, H. Bradley, Minx, Patrick, Warren, Daniel E., Shedlock, Andrew M., Thomson, Robert C., Valenzuela, Nicole, Abramyan, John · 59 authors totalLearning Large-Scale Conditional Random Fields
Figshare · DOI 10.1184/r1/6720377.v1 · 7 citations · Source: semantic-scholar+openalexConditional Random Fields (CRFs) [Lafferty et al., 2001] can offer computational and statistical advantages over generative models, yet traditional CRF parameter and structure learning methods are often too expensive to scale up to large problems. This thesis develops methods capable of learning CRFs for much larger problems. We do so by decomposing learning problems into smaller, simpler subproblems. These decompositions allow us to trade off sample complexity, computational complexity, and potential for parallelization, and we can often optimize these trade-offs in model- or data-specific ways. The resulting methods are theoretically motivated, are often accompanied by strong guarantees, and are effective and highly scalable in practice. In the first part of our work, we develop core methods for CRF parameter and structure learning. For parameter learning, we analyze several methods and produce PAC learnability results for certain classes of CRFs. Structured composite likelihood estimation proves particularly successful in both theory and practice, and our results offer guidance for optimizing estimator structure. For structure learning, we develop a maximum-weight spanning tree-based method which outperforms other methods for recovering tree CRFs. In the second part of our work, we take advantage of the growing availability of parallel platforms to speed up regression, a key component of our CRF learning methods. Our Shotgun algorithm for parallel regression can achieve near-linear speedups, and extensive experiments show it to be one of the fastest methods for sparse regression.
Joseph Bradley, Bradley, Joseph K. · 2 authors total