Computer Science > Computation and Language

arXiv:2405.01481 (cs)

[Submitted on 2 May 2024 (v1), last revised 3 Sep 2024 (this version, v2)]

Title:NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment

Authors:Gerald Shen, Zhilin Wang, Olivier Delalleau, Jiaqi Zeng, Yi Dong, Daniel Egert, Shengyang Sun, Jimmy Zhang, Sahil Jain, Ali Taghibakhshi, Markel Sanz Ausin, Ashwath Aithal, Oleksii Kuchaiev

View PDF HTML (experimental)

Abstract:Aligning Large Language Models (LLMs) with human values and preferences is essential for making them helpful and safe. However, building efficient tools to perform alignment can be challenging, especially for the largest and most competent LLMs which often contain tens or hundreds of billions of parameters. We create NeMo-Aligner, a toolkit for model alignment that can efficiently scale to a thousand GPUs for training the largest open-source LLMs such as Nemotron 4 340B and Llama 3.1 405B. NeMo-Aligner comes with highly optimized and scalable implementations for major paradigms of model alignment such as: Reinforcement Learning from Human Feedback (RLHF), Direct Preference Optimization (DPO), SteerLM, and Self-Play Fine-Tuning (SPIN). Additionally, our toolkit supports running most of the alignment techniques in a Parameter Efficient Fine-Tuning (PEFT) setting. NeMo-Aligner is designed for extensibility, allowing support for other alignment techniques with minimal effort. It is open-sourced with Apache 2.0 License and we invite community contributions at this https URL

Comments:	16 pages, 4 figures, Accepted to COLM 2024
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2405.01481 [cs.CL]
	(or arXiv:2405.01481v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2405.01481

Submission history

From: Zhilin Wang [view email]
[v1] Thu, 2 May 2024 17:13:40 UTC (741 KB)
[v2] Tue, 3 Sep 2024 05:47:42 UTC (753 KB)

Computer Science > Computation and Language

Title:NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators