Data and methods for identifying artificial intelligence-related patents

This paper evaluates existing approaches to identifying artificial intelligence (AI)-related patents and introduces a novel, scalable framework to improve classification performance. Motivated by the growing reliance on patent data in innovation research, we assess widely used methods and document their substantial performance limitations. To address these challenges, we develop a CPC-informed label-refinement framework inspired by positive-unlabeled (PU) learning to construct high-quality training data. Using this refined dataset, we train a range of machine learning, deep learning, and transformer-based models. Our models substantially outperform existing approaches, including those employed by the USPTO. We further demonstrate the empirical value of more accurate AI patent identification through two applications. First, we find that the release of ChatGPT increased the market valuation of AI patents. Second, we show that firms increased the allocation of innovative effort toward AI technologies following its release. To facilitate future research, we release our training data, source code, and a novel dataset of patent-level predictions (AIPat), which will be continuously updated to reflect the evolving nature of AI innovation.

Tianjun Wu (Nanjing University), 
Chao Min (Nanjing University), 
Waverly W. Ding (University of Maryland), 
Guolong Wang (University of International Business and Economics), 
Kunpeng Zhang (University of Maryland)

Research Policy
  • Kunpeng Zhang
  • Waverly Ding
  • Decision, Operations and Information Technologies
  • Corporate strategy and global competitiveness
  • Creativity, innovation and organizational change
  • Policy innovation and strategic foresight
  • Corporate and consumer activism
  • Artificial intelligence (AI)
  • Generative machine learning, data analytics, data fusion, and personalization
  • Information Technology
    Back to Top