Descarga DALE - Descarga del código fuente DALE

DALE

Otro código fuente

1.0.0

Descargar

DALE: Aumento de datos generativos para PNL legal de bajos recursos

Implementación del documento EMNLP 2023: DALE: Generative Data Augmentation for Low-Resource Legal PNL.

Metodología propuesta

El modelo grande de Bart preentrenado de DALE se puede encontrar aquí. Los datos previos al entrenamiento se pueden encontrar aquí.

Pasos:

Instale dependencias usando:
```
 pip install -r requirements.txt
```

Ejecute los archivos requeridos
Para enmascaramiento de PMI:

 cd pmi/
sh pmi.sh <config_name> <dataset_path> <output_path> <n_gram_value> <pmi_cut_off>
sh pmi.sh unfair_tos ./unfair_tos ./output_path 3 95

Para preentrenamiento de BART:

 cd bart_pretrain/
python pretrain.py --ckpt_path ./ckpt_path 
                  --dataset_path ./dataset_path> 
                  --max_input_length 1024 
                  --max_target_length 1024 
                  --batch_size 4
                  --num_train_epochs 10
                  --logging_steps 100
                  --save_steps 1000
                  --output_dir ./output_path

Para generación BART:

 cd bart_generation/
python bart_ctx_augs.py --dataset_name "scotus" 
                  --path ./dataset_path 
                  --dest_path ./dest_path
                  --n_augs 5 
                  --batch_size 4
                  --model_path ./model_path

bart_ctx_augs.py -> BART generation for multi-class data generation.
bart_ctx_augs_multi.py -> BART generation for multi-label data generation.
bart_ctx_augs_ch.py -> BART generation for casehold dataset.

Información de citas y entradas de BibTeX

Si encuentra útil nuestro documento/código/demostración, cite nuestro documento:

 @inproceedings{ghosh-etal-2023-dale,
    title = "DALE: Generative Data Augmentation for Low-Resource Legal NLP",
    author = "Sreyan Ghosh  and
      Chandra Kiran Evuru  and
      Sonal Kumar  and
      S Ramaneswaran and
      S Sakshi and
      Utkarsh Tyagi and
      Dinesh Manocha",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Sentosa, Singapore",
    abstract = "We present DALE, a novel and effective generative Data Augmentation framework for lowresource LEgal NLP. DALE addresses the challenges existing frameworks pose in generating effective data augmentations of legal documents - legal language, with its specialized vocabulary and complex semantics, morphology, and syntax, does not benefit from data augmentations that merely rephrase the source sentence. To address this, DALE, built on an EncoderDecoder Language Model, is pre-trained on a novel unsupervised text denoising objective based on selective masking - our masking strategy exploits the domain-specific language characteristics of templatized legal documents to mask collocated spans of text. Denoising these spans help DALE acquire knowledge about legal concepts, principles, and language usage. Consequently, it develops the ability to generate coherent and diverse augmentations with novel contexts. Finally, DALE performs conditional generation to generate synthetic augmentations for low-resource Legal NLP tasks. We demonstrate the effectiveness of DALE on 13 datasets spanning 6 tasks and 4 low-resource settings. DALE outperforms all our baselines, including LLMs, qualitatively and quantitatively, with improvements of 1%-50%."
}

Expandir

Información adicional

Versión 1.0.0
Tipo Otro código fuente
Fecha de actualización 2024-12-05
tamaño 749.15KB
Proviene de Github

Aplicaciones relacionadas

waymo open dataset

2024-11-18
SmartTube

2024-12-14
Sunamu

2024-12-14
MySchedule.py

2024-12-15
viptools for eslam

2024-12-15
VITAident

2024-12-15

Recomendado para ti

chat.petals.dev

Otro código fuente

1.0.0
GPT Prompt Templates

Otro código fuente

1.0.0
GPTyped

Otro código fuente

GPTyped 1.0.5
waymo open dataset

Otro código fuente

December 2023 Update
SmartTube

Otro código fuente

24.71 Stable
Sunamu

Otro código fuente

Release 2.2.0
waymo open dataset

Otro código fuente

December 2023 Update
wp functions

Otras categorias

1.0.0
termwind

Otras categorias

v2.3.0

Información relacionada Todo