llava Skill
Vision-language chat: VQA, captioning, image dialogue. Published by NousResearch in hermes-agent.
Decision snapshot
Is this a fit?
Design and media, Research, Includes SKILL.md, Reusable instructions
Compatibility not yet detected.
Permission behavior not yet detected.
Copy skill directory
21 days ago · MIT license
No specific cautions were detected. Review the source and requested permissions before installing.
What is llava Skill?
Vision-language chat: VQA, captioning, image dialogue. Published by NousResearch in hermes-agent. This profile combines repository metadata with install, compatibility, and usage signals so developers can quickly decide whether it fits their agent workflow before opening the source repository.
Automated repository signals based on public metadata such as recency, license, installation evidence, and adoption. These are not a security audit or endorsement. See how SkillIndex evaluates profiles.
Key capabilities
- Includes SKILL.md support
- Reusable instructions support
- Design and media
- Research
- Design and media use cases
- Research use cases
Declared skill metadata
- Declared author: Orchestra Research
- Declared license: MIT
- Source file: optional-skills/mlops/llava/SKILL.md
These fields retain source and confidence evidence from the indexed SKILL.md.
Compatibility and setup
- Install or run with Copy skill directory
When to use llava Skill
- Use it for design and media.
- Use it for research.
Built with
Editorial notes
Source
- Creator: NousResearch
- Repository: NousResearch/hermes-agent
- Skill file: optional-skills/mlops/llava/SKILL.md
What it does
Vision-language chat: VQA, captioning, image dialogue.
Skill instructions
LLaVA - Large Language and Vision Assistant Open-source vision-language model for conversational image understanding. When to use LLaVA Use when: - Building vision-language chatbots - Visual question answering (VQA) - Image description and captioning - Multi-turn image conversations - Visual instruction following - Document understanding with images Metrics: - 23,000+ GitHub stars - GPT-4V level capabilities (targeted) - Apache 2.0 License - Multiple model sizes (7B-34B params) Use alternatives instead: - GPT-4V: Highest quality, API-based - CLIP: Simple zero-shot classification - BLIP-2: Better for captioning only - Flamingo: Research, not open-source Quick start Installation bash Clone repository git clone https://github.com/haotian-liu/LLaVA cd LLaVA Install pip install -e . Basic usage python from llava.model.builder import loadpretrainedmodel from llava.mmutils import getmodelnamefrompath, processimages, tokenizerimagetoken from llava.constants import IMAGETOKENINDEX, DEFAULTIMAGE
Verified compatibility and discovery
Frequently asked questions
What is llava?
llava is a open-source AI agent skill with Copy skill directory. Vision-language chat: VQA, captioning, image dialogue.
Who is llava best for?
llava is best for reusing agent instructions, scripts, and references, design and media workflows, research workflows.
How do I install llava?
Install or run llava using Copy skill directory. Check llava for the latest setup command.
Is llava actively maintained?
llava may need a closer maintenance check before production use.
Project health auto-fetched from the source repository.
Maintain this resource?
Review this source-backed profile, send a correction with evidence, or link to it from your documentation. Claims verify your relationship to the project; profile facts still require source evidence and editorial review.