CLIP Vision
Clip Vision H in ComfyUI
22 workflows use this model
What is Clip Vision H?
Clip Vision H is a CLIP Vision encoder that converts images into embeddings for conditioning or style transfer. You can run it locally in ComfyUI with full control over every parameter, or access it through Comfy Cloud. ComfyUI's node-based workflow editor lets you connect Clip Vision H with ControlNets, LoRAs, upscalers, and custom nodes to build any pipeline you need. There are 22 community workflow templates using Clip Vision H on Comfy Workflows, ready to load and customize.
Read the full tutorial →Frequently Asked Questions
Clip Vision H is a CLIP Vision encoder that converts images into embeddings for conditioning or style transfer. You can run it locally in ComfyUI with full control over every parameter, or access it through Comfy Cloud. ComfyUI's node-based workflow editor lets you connect Clip Vision H with ControlNets, LoRAs, upscalers, and custom nodes to build any pipeline you need. There are 22 community workflow templates using Clip Vision H on Comfy Workflows, ready to load and customize.
Follow the step-by-step tutorial at https://docs.comfy.org/tutorials/video/wan/wan2-2-animate. You can also load any of the 22 community workflow templates that use Clip Vision H directly in ComfyUI.
There are 22 community workflow templates that use Clip Vision H on Comfy Workflows. Each template is ready to run in ComfyUI and can be customized to suit your project.
ComfyUI is free and open source. Clip Vision H weights are available to download from Hugging Face. You only pay for compute when running on Comfy Cloud; local inference on your own hardware is always free.