<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>CV on 南巷</title><link>https://nx.wmlab.top/tags/cv/</link><description>Recent content in CV on 南巷</description><generator>Hugo -- gohugo.io</generator><language>zh-CN</language><managingEditor>liuwenhao1968@163.com (南巷)</managingEditor><webMaster>liuwenhao1968@163.com (南巷)</webMaster><copyright>© 2026 南巷</copyright><lastBuildDate>Mon, 15 Jun 2026 02:14:39 +0800</lastBuildDate><atom:link href="https://nx.wmlab.top/tags/cv/index.xml" rel="self" type="application/rss+xml"/><item><title>ViT</title><link>https://nx.wmlab.top/blog/vit/</link><pubDate>Mon, 15 Jun 2026 02:14:39 +0800</pubDate><author>liuwenhao1968@163.com (南巷)</author><guid>https://nx.wmlab.top/blog/vit/</guid><description>下面详细分析一下ViT，这个开启了cv的新时代。在cnn处理不好的地方，在vit却是可以很好的处理。</description></item><item><title>SegFormer</title><link>https://nx.wmlab.top/blog/segformer/</link><pubDate>Mon, 15 Jun 2026 02:14:35 +0800</pubDate><author>liuwenhao1968@163.com (南巷)</author><guid>https://nx.wmlab.top/blog/segformer/</guid><description>提出一个层级（多尺度）且不依赖位置编码的 Transformer 编码器（MiT），可以直接输出多尺度特征（1/4、1/8、1/16、1/32），避免了 ViT 在不同测试分辨率下必须插值位置编码带来的性能下降。</description></item><item><title>Swin Transformer</title><link>https://nx.wmlab.top/blog/swin-transformer/</link><pubDate>Mon, 15 Jun 2026 02:14:32 +0800</pubDate><author>liuwenhao1968@163.com (南巷)</author><guid>https://nx.wmlab.top/blog/swin-transformer/</guid><description>Patch 分割和线性嵌入 \(Patch Partition + Linear Embedding\)</description></item><item><title>MaskFormer</title><link>https://nx.wmlab.top/blog/maskformer/</link><pubDate>Mon, 15 Jun 2026 02:14:28 +0800</pubDate><author>liuwenhao1968@163.com (南巷)</author><guid>https://nx.wmlab.top/blog/maskformer/</guid><description>具体表现：</description></item><item><title>Mask2Former</title><link>https://nx.wmlab.top/blog/mask2former/</link><pubDate>Mon, 15 Jun 2026 02:14:20 +0800</pubDate><author>liuwenhao1968@163.com (南巷)</author><guid>https://nx.wmlab.top/blog/mask2former/</guid><description>MaskFormer 让分割任务变成 “预测一组掩码 + 类别”； Mask2Former 让模型 “只看自己负责的掩码区域，更聪明地预测掩码”。</description></item><item><title>DINO</title><link>https://nx.wmlab.top/blog/dino/</link><pubDate>Mon, 15 Jun 2026 02:14:13 +0800</pubDate><author>liuwenhao1968@163.com (南巷)</author><guid>https://nx.wmlab.top/blog/dino/</guid><description>在 DINO 出现之前，自监督视觉表示学习主要有两类主流方法：</description></item><item><title>SAM</title><link>https://nx.wmlab.top/blog/sam/</link><pubDate>Mon, 15 Jun 2026 02:14:12 +0800</pubDate><author>liuwenhao1968@163.com (南巷)</author><guid>https://nx.wmlab.top/blog/sam/</guid><description>sam设计的出发点：</description></item><item><title>DETR</title><link>https://nx.wmlab.top/blog/detr/</link><pubDate>Mon, 15 Jun 2026 02:14:04 +0800</pubDate><author>liuwenhao1968@163.com (南巷)</author><guid>https://nx.wmlab.top/blog/detr/</guid><description>在做DETR之前，目标检测都是下面两种任务方式：</description></item><item><title>Deformable DETR</title><link>https://nx.wmlab.top/blog/deformable-detr/</link><pubDate>Mon, 15 Jun 2026 02:14:02 +0800</pubDate><author>liuwenhao1968@163.com (南巷)</author><guid>https://nx.wmlab.top/blog/deformable-detr/</guid><description>原来大DETR有几个缺点：</description></item></channel></rss>