Spaces:

prithivMLmods
/

Qwen-Image-Diffusion

Running on Zero

App Files Files Community

prithivMLmods commited on 18 days ago

Commit

0ed0cc1

verified ·

1 Parent(s): 68eb731

Update app.py

Browse files

Files changed (1) hide show

app.py +5 -3

app.py CHANGED Viewed

@@ -345,10 +345,12 @@ with gr.Blocks(css=css, theme="bethecloud/storj_theme") as demo:
             ],
                                     label="Select Model",
                                     value="Camel-Doc-OCR-080125(v2)")
             gr.Markdown("**Model Info 💻** | [Report Bug](https://huggingface.co/spaces/prithivMLmods/Multimodal-OCR-Outpost/discussions)")
     # Define the submit button actions
     image_submit.click(fn=generate_image,
                        inputs=[

             ],
                                     label="Select Model",
                                     value="Camel-Doc-OCR-080125(v2)")
             gr.Markdown("**Model Info 💻** | [Report Bug](https://huggingface.co/spaces/prithivMLmods/Multimodal-OCR-Outpost/discussions)")
+            gr.Markdown("> Camel-Doc-OCR-080125 is a specialized vision-language model, fine-tuned from Qwen2.5-VL-7B-Instruct, and excels at document retrieval, content extraction, and analysis recognition for both structured and unstructured digital documents. OCRFlux-3B is a 3B-parameter vision-language model optimized for high-quality OCR on PDFs and images, excelling in converting documents to clean Markdown text and supporting features like cross-page table/paragraph merging.")
+            gr.Markdown("> Both ViGoRL-MCTS-SFT-3b-Spatial and 7b-Spatial are vision-language models that use multi-turn visually grounded reinforcement learning for precise spatial reasoning and visual grounding, with the 3b and 7b variants differing mainly in their architectural size for fine-grained visual tasks.")
+            gr.Markdown("> Owlet-safety-3b-1 is tailored for multi-label safety event detection in video, such as identifying fire, smoke, fall, and theft, and is useful for surveillance and incident detection scenarios. MonkeyOCR-pro-1.2B is a lightweight yet powerful model for document parsing using a structure-recognition-relation triplet method, delivering state-of-the-art OCR results.")
     # Define the submit button actions
     image_submit.click(fn=generate_image,
                        inputs=[