Salesforce
/

LLaMA-3-8B-SFR-SFT-R

Text Generation

text-generation-inference

Inference Endpoints

Model card Files Files and versions Community

Edit model card

LLaMA-3-8B-SFR-SFT-R

This is the SFT model for Salesforce/SFR-Iterative-DPO-LLaMA-3-8B-R.

Model Releases

Citation

Please cite our techical report if you find our model is useful for your research or product.

@misc{dong2024rlhf,
      title={RLHF Workflow: From Reward Modeling to Online RLHF}, 
      author={Hanze Dong and Wei Xiong and Bo Pang and Haoxiang Wang and Han Zhao and Yingbo Zhou and Nan Jiang and Doyen Sahoo and Caiming Xiong and Tong Zhang},
      year={2024},
      eprint={2405.07863},
      archivePrefix={arXiv},
      primaryClass={cs.LG}
}

Downloads last month: 14

Safetensors

Model size

8.03B params

Tensor type

BF16

·

Inference Examples

Text Generation

This model does not have enough activity to be deployed to Inference API (serverless) yet. Increase its social visibility and check back later, or deploy to Inference Endpoints (dedicated) instead.

Model tree for Salesforce/LLaMA-3-8B-SFR-SFT-R

Quantizations

1 model