Preference & Alignment (DPO/RLHF)RLHFCommercial OK
hh-rlhf
by Anthropic
31.9Kdownloads
1.7Klikes
Description
Dataset Card for HH-RLHF
Dataset Summary
This repository provides access to two different kinds of data:
Human preference data about helpfulness and harmlessness from Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. These data are meant to train preference (or reward) models for subsequent RLHF training. These data are not meant for supervised training of dialogue agents. Training dialogue agents on these data is likely to lead… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/hh-rlhf.
What can I do with this?
Tags
license:mitsize_categories:100K<n<1Mformat:jsonmodality:textlibrary:datasetslibrary:dasklibrary:mlcroissantlibrary:polarsarxiv:2204.05862region:ushuman-feedback