CAIL2019-SCM: A Dataset of Similar Case Matching in Legal Domain

TLDR

The dataset and additional information are available on GitHub. The paper introduces the CAIL2019‑SCM dataset for similar case matching in Chinese law. The dataset comprises 8,964 triplets of Supreme People’s Court cases, where participants must identify the most similar pair within each triplet. The competition attracted 711 teams, with the top score of 71.88, and several baseline methods were provided.

Abstract

In this paper, we introduce CAIL2019-SCM, Chinese AI and Law 2019 Similar Case Matching dataset. CAIL2019-SCM contains 8,964 triplets of cases published by the Supreme People's Court of China. CAIL2019-SCM focuses on detecting similar cases, and the participants are required to check which two cases are more similar in the triplets. There are 711 teams who participated in this year's competition, and the best team has reached a score of 71.88. We have also implemented several baselines to help researchers better understand this task. The dataset and more details can be found from https://github.com/china-ai-law-challenge/CAIL2019/tree/master/scm.

References

Page 1

	Year	Citations

Page 1