Is the actual context length 32768?

#1
by MuAlphaOmegaEpsilon - opened

I tried to launch > llama-server LFM2.5-VL-3B-Q4_0.gguf -c 131072just for testing and it reported the following:

0.02.592.680 W llama_context: n_ctx_seq (131072) > n_ctx_train (128000) -- possible training context overflow
0.03.071.738 W srv    load_model: the slot context (131072) exceeds the training context of the model (128000) - capping

Are you sure the context length is limited to 32,768 as reported in the Base model card?
I'm trying to figure out if this model can be used in place of the standard LFM2.5-2.6B with the additional vision support on top or if I have to make some compromises on the context length.

MuAlphaOmegaEpsilon changed discussion title from Is the actual context length to Is the actual context length 32768?

+1 Would love to know the same

32,768 is the maximum context length seen during vision training. Different phases during text-only pre-training see longer contexts; but the model has never seen examples with an image and a longer context. So you should run it with -c 32768, and we'll try to fix the metadata in the GGUFs.

I'm trying to figure out if this model can be used in place of the standard LFM2.5-2.6B with the additional vision support on top or if I have to make some compromises on the context length.

Unfortunately, I think LFM2.5-VL-3B is not nearly as agentic as the 2.6B. We will try to improve the VL models' agentic capabilities in future releases, but if you absolutely need multimodal, agentic capabilities on edge right now, I've seen some demos that use the VL model to describe/perform OCR on images, then pass the input to the 2.6B. Might be worth hacking around and seeing if it works.

If you have particular use cases that you would like to address, please let me know! You can find my email on my website.

Thank you for the comment on this <3
I'll be experimenting later how much I can push the 2.6B model for agentic codebase discovery, and general web-scraping and document-managing tasks, among other things and experiments regarding context management - considering my custom harness in the mix. VL combination as a separately hosted rest listener will probably be the route I'll go then, but a multimodal version of 2.5 would always be welcome <3

Sign up or log in to comment