Sentence Embeddings and High-Speed Similarity Search for Fast Computer Assisted Annotation of Legal Documents

AUTHORS> Hannes Westermann; Jaromír Šavelka; Vern R. Walker; Kevin D. Ashley; Karim Benyekhlef

2020 · 2020-12-24 · Frontiers in Artificial Intelligence and Applications; Legal Knowledge and Information Systems

Read at the source ↗DOI 10.3233/faia200860

ABSTRACT

Abstract

Human-performed annotation of sentences in legal documents is an important prerequisite to many machine learning based systems supporting legal tasks. Typically, the annotation is done sequentially, sentence by sentence, which is often time consuming and, hence, expensive. In this paper, we introduce a proof-of-concept system for annotating sentences "laterally." The approach is based on the observation that sentences that are similar in meaning often have the same label in terms of a particular type system. We use this observation in allowing annotators to quickly view and annotate sentences that are semantically similar to a given sentence, across an entire corpus of documents. Here, we present the interface of the system and empirically evaluate the approach. The experiments show that lateral annotation has the potential to make the annotation process quicker and more consistent.

CC BY-NC 4.0 · Source ↗

AUTHOR KEYWORDS> Annotation; Language Models; Sentence Embeddings; Approximate Nearest Neighbour; Interactive Machine Learning

VERIFICATION

Full text checked · 2026-10-07