Favicon of llava

llava Skill

AI Agent SkillPythonOpen source

Vision-language chat: VQA, captioning, image dialogue. Published by NousResearch in hermes-agent.

Decision snapshot

Is this a fit?

Best for

Design and media, Research, Includes SKILL.md, Reusable instructions

Works with

Compatibility not yet detected.

Access

Permission behavior not yet detected.

Setup

Copy skill directory

Project health

21 days ago · MIT license

Considerations

No specific cautions were detected. Review the source and requested permissions before installing.

What is llava Skill?

Vision-language chat: VQA, captioning, image dialogue. Published by NousResearch in hermes-agent. This profile combines repository metadata with install, compatibility, and usage signals so developers can quickly decide whether it fits their agent workflow before opening the source repository.

Trust signal
95/100
Maintenance signal
90/100
Adoption signal
100/100

Automated repository signals based on public metadata such as recency, license, installation evidence, and adoption. These are not a security audit or endorsement. See how SkillIndex evaluates profiles.

Key capabilities

  • Includes SKILL.md support
  • Reusable instructions support
  • Design and media
  • Research
  • Design and media use cases
  • Research use cases

Declared skill metadata

  • Declared author: Orchestra Research
  • Declared license: MIT
  • Source file: optional-skills/mlops/llava/SKILL.md

These fields retain source and confidence evidence from the indexed SKILL.md.

Compatibility and setup

Copy skill directory
  • Install or run with Copy skill directory

When to use llava Skill

  • Use it for design and media.
  • Use it for research.

Built with

PythonCopy skill directory

Editorial notes

Source

  • Creator: NousResearch
  • Repository: NousResearch/hermes-agent
  • Skill file: optional-skills/mlops/llava/SKILL.md

What it does

Vision-language chat: VQA, captioning, image dialogue.

Skill instructions

LLaVA - Large Language and Vision Assistant Open-source vision-language model for conversational image understanding. When to use LLaVA Use when: - Building vision-language chatbots - Visual question answering (VQA) - Image description and captioning - Multi-turn image conversations - Visual instruction following - Document understanding with images Metrics: - 23,000+ GitHub stars - GPT-4V level capabilities (targeted) - Apache 2.0 License - Multiple model sizes (7B-34B params) Use alternatives instead: - GPT-4V: Highest quality, API-based - CLIP: Simple zero-shot classification - BLIP-2: Better for captioning only - Flamingo: Research, not open-source Quick start Installation bash Clone repository git clone https://github.com/haotian-liu/LLaVA cd LLaVA Install pip install -e . Basic usage python from llava.model.builder import loadpretrainedmodel from llava.mmutils import getmodelnamefrompath, processimages, tokenizerimagetoken from llava.constants import IMAGETOKENINDEX, DEFAULTIMAGE

Verified compatibility and discovery

Frequently asked questions

What is llava?

llava is a open-source AI agent skill with Copy skill directory. Vision-language chat: VQA, captioning, image dialogue.

Who is llava best for?

llava is best for reusing agent instructions, scripts, and references, design and media workflows, research workflows.

How do I install llava?

Install or run llava using Copy skill directory. Check llava for the latest setup command.

Is llava actively maintained?

llava may need a closer maintenance check before production use.

Share:

Stars
232,138
Forks
46,253
Last commit
21 days ago
Last verified
Aug 18, 2026
Metadata fetched
Aug 18, 2026
Repository age
1 year
License
MIT

Project health auto-fetched from the source repository.

Maintain this resource?

Review this source-backed profile, send a correction with evidence, or link to it from your documentation. Claims verify your relationship to the project; profile facts still require source evidence and editorial review.

Alternatives to llava