---
title: "Skill issue: stop deploying vision language models, use them with Skills to build e2e vision apps on edge"
category: "talks"
date: "2026-06-29"
time: "11:40am-12:00pm"
track: "Vision & OCR"
room: "Track 2"
speakers: ["Merve Noyan"]
sourceLabels: ["Official conference schedule", "Public YouTube metadata"]
scheduleTrack: "Vision & OCR"
scheduleRoom: "Track 2"
scheduleLabels: ["Vision & OCR", "Track 2", "sponsor", "confirmed"]
---
# Skill issue: stop deploying vision language models, use them with Skills to build e2e vision apps on edge

## Conference Context
- Date/time: 2026-06-29 · 11:40am-12:00pm
- Track/room: Vision & OCR · Track 2
- Speaker(s): Merve Noyan
- Session type/status: sponsor · confirmed

- Track: Vision & OCR
- Room: Track 2
- Session type: sponsor
- Status: confirmed

## Session Description
With the boom of vision language models barrier of entry to build vision apps are much lower so developers tend to use them right away. However, these models are very large and inefficient in production. In this talk, I will go through combining vision language models with Skills to build end-to-end vision apps from training to deployment using HF Skills, on top of showing the state-of-the-art in small computer vision/multimodal models.

## Media Evidence
[Self-Training Agents: Hermes Agent, HF Traces, Skills, MCP & Finetuning  — Merve Noyan, Hugging Face](https://www.youtube.com/watch?v=OV56RddyFuU) (speaker-match related prior/adjacent AI Engineer video; captions: English auto-captions).

- Source video: `youtube-OV56RddyFuU`
- Slide deck: [[youtube-OV56RddyFuU-dense-slides|Dense Slides: Self-Training Agents: Hermes Agent, HF Traces, Skills, MCP & Finetuning  — Merve Noyan, Hugging Face]] — 17 visible slide image(s); 17 HTML recreation(s).
![[assets/dense-slides/OV56RddyFuU/slide-001.jpg]]
![[assets/dense-slides/OV56RddyFuU/slide-002.jpg]]
![[assets/dense-slides/OV56RddyFuU/slide-003.jpg]]
- Additional slide evidence: [[youtube-OV56RddyFuU-slides|Slides: Self-Training Agents: Hermes Agent, HF Traces, Skills, MCP & Finetuning  — Merve Noyan, Hugging Face]], [[youtube-OV56RddyFuU-reconstructed-slides|Reconstructed Slides: Self-Training Agents: Hermes Agent, HF Traces, Skills, MCP & Finetuning  — Merve Noyan, Hugging Face]]
- Slide-derived themes for `youtube-OV56RddyFuU`: models, does, matter, absolute, control, over, cost, reduction.

## Evidence Graph
This evidence graph is generated from currently linked source material: official schedule text, related video pages, cached transcripts, visible slide text, dense/reconstructed slide pages, and AI slide-classification audits.

### Media Signals
- `youtube-OV56RddyFuU` — 6 slide-derived text signals
- Slide-derived themes for `youtube-OV56RddyFuU`: models, does, matter, absolute, control, over, cost, reduction.
- Evidence links for `youtube-OV56RddyFuU`: [[youtube-OV56RddyFuU]], [[youtube-OV56RddyFuU-slides]], [[youtube-OV56RddyFuU-dense-slides]], [[youtube-OV56RddyFuU-reconstructed-slides]]

### Agent Reading Notes
Use these signals to refine the synopsis, topic links, people/company context, and method notes. If a source is a related external video rather than an exact official recording, keep it framed as supporting evidence.

## Transcript Status
Related video transcript availability: English auto-captions. Treat this as supporting context, not a recording of this exact scheduled session unless later confirmed. Not fetched yet.

## People
- [[merve-noyan]]

## Supporting Slides
- [[youtube-OV56RddyFuU-slides]] — extracted from the related public AI Engineer video.

## Slide Evidence
- Slide-only cropped deck: [[youtube-OV56RddyFuU-dense-slides]] (20 viable slide images).
- Related slide/OCR pages:
- [[youtube-OV56RddyFuU-dense-slides]]
- [[youtube-OV56RddyFuU-reconstructed-slides]]
- [[youtube-OV56RddyFuU-slides]]
- Slide-derived terms: `model`, `skills`, `models`, `gguf`, `hermes`, `community`, `datasets`, `text`, `local`, `llama.cpp`, `inference`, `image`, `search`, `spaces`, `jobs`, `setup`, `claude`, `some`

## Synthesis
### Synthesized Breakdown
# Skill issue: stop deploying vision language models, use them with Skills to build e2e vision apps on edge ## Conference Context - Date/time: 2026-06-29 · 11:40am-12:00pm - Track/room: Vision & OCR · Track 2 - Speaker(s): Merve Noyan - Session type/status: sponsor · confirmed - Track: Vision & OCR - Room: Track 2 - Session type: sponsor - Status: confirmed ## Session Description With the boom of vision language models barrier of entry to build vision apps are much lower so developers tend to use them right away. However, these models are very large and inefficient in production. In this talk, I will go through combining vision language models with Skills to build end-to-end vision apps from training to deployment using HF Skills, on top of showing the state-of-the-art in small computer vision/multimodal models. ## Media Evidence [Self-Training Agents: Hermes Agent, HF Traces, Skills, MCP & Finetuning — Merve Noyan, Hugging Face](https://www.youtube.com/watch?v=OV56RddyFuU) (speaker-match related prior/adjacent AI Engineer video; captions: English auto-captions).

### Speaker And Company Context
- [[merve-noyan|Merve Noyan]] — MLE at [[hugging-face|Hugging Face]].

### Topics Covered
- [[agent-security]]
- [[agentic-search]]
- [[coding-agents]]
- [[mcp]]

### Derived Links And Source Material
- [[youtube-OV56RddyFuU]] — related YouTube source page.
- [[youtube-OV56RddyFuU-slides]] — slide evidence.
- [[youtube-OV56RddyFuU-reconstructed-slides]] — slide evidence.
- [[youtube-OV56RddyFuU-dense-slides]] — slide evidence.

### Novel Concepts / Clever Methods
- No highlighted novel concept has been detected yet.

### Evidence Boundary
This synthesis is based on the official schedule and linked source pages. It should be revisited when exact session recordings or transcript-backed secondary sources are available.
