2018 · 25 citations · 8 references
Artificial IntelligenceEngineeringMachine LearningLawRisk AnalysisCorpus LinguisticsText MiningNatural Language ProcessingInformation RetrievalData ScienceData MiningBetter Decision SupportComputational LinguisticsDocument ClassificationParagraph VectorLegal Information RetrievalMachine Learning CapabilityNlp TaskKnowledge DiscoveryTerminology ExtractionDecision Support SystemsAnalyse RiskAutomated Decision-makingInformation ExtractionIntelligent Decision Support SystemBusinessKeyword ExtractionIntelligent Decision Making
Assessing risk for voluminous legal documents such as request for proposal, contracts is tedious and error prone. We have developed "risk-o-meter", a framework, based on machine learning and natural language processing to review and assess risks of any legal document. Our framework uses Paragraph Vector, an unsupervised model to generate vector representation of text. This enables the framework to learn contextual relations of legal terms and generate sensible context aware embedding. The framework then feeds the vector space into a supervised classification algorithm to predict whether a paragraph belongs to a pre-defined risk category or not. The framework thus extracts risk prone paragraphs. This technique efficiently overcomes the limitations of keyword based search. We have achieved an accuracy of 91% for the risk category having the largest training dataset. This framework will help organizations optimize effort to identify risk from large document base with minimal human intervention and thus will help to have risk mitigated sustainable growth. Its machine learning capability makes it scalable to uncover relevant information from any type of document apart from legal documents, provided the library is pre-populated and rich.
8
Efficient Estimation of Word Representations in Vector Space
Tomáš Mikolov, Kai Chen, Greg S. Corrado · arXiv (Cornell University) · 2013 · 18.1K citations · Full text
Efficient Estimation of Word Representations in Vector Space
Tomáš Mikolov, Kai Chen, Greg S. Corrado et al. · arXiv (Cornell University) · 2013 · 11.7K citations · Full text
A unified architecture for natural language processing
Ronan Collobert, Jason Weston · 2008 · 5.2K citations
Engineering, Machine Learning, Cross-lingual Representation +19
Distributed Representations of Sentences and Documents
Quoc V. Le, Tomáš Mikolov · arXiv (Cornell University) · 2014 · 5.1K citations · Full text
Dependency-Based Word Embeddings
Omer Levy, Yoav Goldberg · 2014 · 1.1K citations · Full text