Edit on GitHub

model = AutoAdapterModel.from_pretrained("gpt2")
config = AdapterConfig.load("houlsby", non_linearity="swish", reduction_factor=16)
model.load_adapter("nli/rte@ukp", config=config)

Description

Adapter for gpt2 in Houlsby architecture trained on the RTE dataset for 10 epochs with a learning rate of 1e-4.

Properties

Pre-trained model
gpt2
Adapter type
Prediction Head
  Yes
Task
Natural Language Inference
Dataset

Architecture

Name
houlsby
Non-linearity
swish
Reduction factor
16
{
  "ln_after": false,
  "ln_before": false,
  "mh_adapter": true,
  "output_adapter": true,
  "adapter_residual_before_ln": false,
  "non_linearity": "swish",
  "original_ln_after": true,
  "original_ln_before": false,
  "reduction_factor": 16,
  "residual_before_ln": true
}

Author

  Name
Hannah Sterz
  Twitter

Versions

Identifier Comment Score Download
1 DEFAULT

Citations

Architecture
@misc{houlsby2019parameterefficient,
  title={Parameter-Efficient Transfer Learning for NLP},
  author={Neil Houlsby and Andrei Giurgiu and Stanislaw Jastrzebski and Bruna Morrone and Quentin de Laroussilhe and Andrea Gesmundo and Mona Attariyan and Sylvain Gelly},
  year={2019},
  eprint={1902.00751},
  archivePrefix={arXiv},
  primaryClass={cs.LG}
}
Task
@inproceedings{bentivogli2009fifth,
  title={The Fifth PASCAL Recognizing Textual Entailment Challenge.},
  author={Bentivogli, Luisa and Clark, Peter and Dagan, Ido and Giampiccolo, Danilo},
  booktitle={TAC},
  year={2009}
}