Instructions to use fancyfeast/llama-joycaption-beta-one-hf-llava with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use fancyfeast/llama-joycaption-beta-one-hf-llava with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="fancyfeast/llama-joycaption-beta-one-hf-llava") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("fancyfeast/llama-joycaption-beta-one-hf-llava") model = AutoModelForMultimodalLM.from_pretrained("fancyfeast/llama-joycaption-beta-one-hf-llava", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use fancyfeast/llama-joycaption-beta-one-hf-llava with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "fancyfeast/llama-joycaption-beta-one-hf-llava" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "fancyfeast/llama-joycaption-beta-one-hf-llava", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/fancyfeast/llama-joycaption-beta-one-hf-llava
- SGLang
How to use fancyfeast/llama-joycaption-beta-one-hf-llava with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "fancyfeast/llama-joycaption-beta-one-hf-llava" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "fancyfeast/llama-joycaption-beta-one-hf-llava", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "fancyfeast/llama-joycaption-beta-one-hf-llava" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "fancyfeast/llama-joycaption-beta-one-hf-llava", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use fancyfeast/llama-joycaption-beta-one-hf-llava with Docker Model Runner:
docker model run hf.co/fancyfeast/llama-joycaption-beta-one-hf-llava
Anyone have a script to run this with vLLM?
I have been running this on vLLM for months without any issue. Last night RunPod sent me an email that there was an issue with my pod. Which turns out to be they literally wiped it out. Now I am trying to get it setup again, but my scripts I had to run it, are no longer working, with a variety of errors. I am guessing something has changed with vLLM that is causing these issues. Wondering if anyone has a working setup script they could share?
Here is my setup script:
I know that the docker image for vLLM v0.10.2 works, so I would try to install specifically v0.10.2 of vLLM, or if possible use that docker image as the base for your pod (I'm not familiar with what runpod requires).
If that fails let me know what versions of vLLM, torch, and transformers your pod is running with.


