PLN | 🇬🇧 NLP

R0:088528eaae518c6f09835c249f9a8635-The Word is Mightier than the Label: Learning without Pointillistic Labels using Data Programming -

The Word is Mightier than the Label: Learning without Pointillistic Labels using Data Programming

We analyze the math fundamentals behind DP and demonstrate the power of it by applying it on two real-world text classification tasks. Furthermore, we compare DP with pointillistic active and semi-supervised learning techniques traditionally applied in data-sparse settings.

The Word is Mightier than the Label: Learning without Pointillistic Labels using Data Programming Leer más »

TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP

TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP

This paper introduces TextAttack, a Python framework for adversarial attacks, data augmentation, and adversarial training in NLP. TextAttack builds attacks from four components: a goal function, a set of constraints, a transformation, and a search method. TextAttack’s modular design enables researchers to easily construct attacks from combinations of novel and existing components. TextAttack provides implementations of 16 adversarial attacks from the literature and supports a variety of models and datasets, including BERT and other transformers, and all GLUE tasks.

TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP Leer más »

Microsoft NLP Best Practices

Microsoft NLP Best Practices

This repository contains examples and best practices for building NLP systems, provided as Jupyter notebooks and utility functions. The focus of the repository is on state-of-the-art methods and common scenarios that are popular among researchers and practitioners working on problems involving text and language

Microsoft NLP Best Practices Leer más »

Open Research Dataset (CORD-19): Semantic Scholar has partnered with leading research groups to release the COVID-19

Open Research Dataset (CORD-19): Semantic Scholar has partnered with leading research groups to release the COVID-19

The Allen Institute just published the #covid19 open research #dataset. In addition, they are sponsoring a related Kaggle competition. The dataset contains almost 30k scholarly articles related to the virus. The goal is to use #NLP to advance our understanding.

Open Research Dataset (CORD-19): Semantic Scholar has partnered with leading research groups to release the COVID-19 Leer más »

Natural Language Processing Succinctly

Natural Language Processing Succinctly

AI assistants represent a significant frontier for development. But the complexities of such systems pose a significant barrier for developers. In Natural Language Processing Succinctly, author Joseph Booth will guide readers through designing a simple system that can interpret and provide reasonable responses to written English text. With this foundation, readers will be prepared to tackle the greater challenges of natural language development. (Syncfusion).

Natural Language Processing Succinctly Leer más »