Stop Removing Stopwords: How an Inherited Preprocessing Default Distorts Legal Text-as-Data
Researchers have found that removing common words from legal documents during data preprocessing can negatively affect the accuracy of empirical legal studies. Because legal analysis often relies on subtle nuances in phrasing, standard data cleaning practices can inadvertently strip away essential context needed for effective machine learning and linguistic classification.
Covered by 1 source
- AarXiv CS.AI↗Gregory M. Dickinson4d ago