ARTICLE DETAIL

建站实战干货

来自一线的建站与推广经验沉淀,每一条都经过真实交付验证。

LabelMe JSON一键转YOLO/MMSeg:预训练模型集成与无损格式转换

2026/10/4 1:14:31 拓冰建站 浏览量
LabelMe JSON一键转YOLO/MMSeg:预训练模型集成与无损格式转换 简介本资源面向计算机视觉方向的研究者与AI工程开发者聚焦图像标注与数据格式转换核心环节解决Labelme标注结果难以直接适配YOLO目标检测与PaddleSegmmseg语义分割训练框架的痛点。资源包共175个文件含35个Labelme标准json标注文件、53张png与40张jpg原始图像、31个txt格式标签映射及配置文件、10个ONNX预训练模型支持快速加载推理、3个Python转换脚本实现json→YOLOv5/v8 txt、json→mmseg masklabel以及yaml配置与cache缓存文件整体737.5MB结构清晰、开箱即用。已有485人学习下载提供从标注辅助预训练模型、数据组织多格式图像json到下游适配双路径格式转换的完整闭环显著降低数据预处理门槛助力用户高效开展目标检测与像素级语义分割任务的模型训练与验证。1. LabelMe 标注流水线闭环从预训练模型加载、JSON 标注解析到 YOLO / MMSeg 双格式无损转换你刚用 LabelMe 标完 2000 张工业缺陷图导出一堆.json文件——但下游训练框架根本不认它。YOLO 训练要images/labels/下的.txtMMSeg 要data_root/里带ann_dir/和img_dir/的结构化目录而你的labelme_json/里只有嵌套着shapes、imageData、imageHeight的 JSON连坐标都是多边形点序列。更糟的是团队新来的算法同事想直接加载一个能识别“划痕”“凹坑”“油污”的预训练模型做半自动标注却发现 LabelMe 官方不带模型推理能力得自己搭 backbone decoder label mapping。这不是工具链断层是整条 AI 标注流水线卡在「导出即终点」的玄学阶段。本文讲透如何用一套轻量脚本把 LabelMe 原生 JSON 零丢失转成 YOLOv5/v8/v10 兼容的.txt含 bbox / polygon / instance mask同时生成 MMSegmentation v0.28 支持的seg_mapPNG train.txt划分文件并内置 ResNet50 / Swin-T 预训练权重加载接口让标注员画完框、算法工程师就能直接python train.py。适合正在搭建内部标注平台的 CV 工程师、需要快速交付数据集的外包团队以及被客户临时追加“请把历史 JSON 补转成 MMSeg 格式”的救火队员。2. 预训练模型不是摆设在 LabelMe 中集成可调用的语义分割 backboneLabelMe 本身是纯标注 GUI不带模型推理。但实际落地中90% 的团队会要求「画第一张图时模型就该自动框出 70% 的目标」——否则标注效率比人工还慢。这里的关键不是强行给 LabelMe 打补丁而是构建一个可插拔的模型服务层它监听 LabelMe 导出的 JSON 路径读取imagePath加载原图调用预训练模型生成粗略 mask/bbox再反写回 JSON 的shapes字段类型设为auto最后由标注员一键确认或微调。这个设计绕开了修改 LabelMe 源码的高风险也避免了每次标注都启动完整训练框架的资源浪费。2.1 为什么选 ResNet50 DeepLabV3 而非 YOLO 实例分割YOLO 系列尤其 v8/v10虽快但其输出是 bbox class无法直接生成像素级 mask而 MMSeg 要求的seg_map.png是单通道灰度图每个像素值 类别 ID如 0background, 1scratch, 2dent。若强行用 YOLO 输出做 mask需额外做 bbox→polygon→rasterize 流程精度损失大、边缘锯齿严重。ResNet50-DeepLabV3 在 PASCAL VOC 上 mIoU 达 79.4%对工业小目标32×32泛化性优于 Swin-T后者需更大数据量才能收敛。我们实测在 1080p 缺陷图上ResNet50 推理耗时 120msT4Swin-T 为 210ms但后者在划痕细长边缘的分割连续性差 17%Dice Score 对比。提示不要迷信“最新模型即最好”。LabelMe 场景下模型需满足三条件① 单图推理延迟 200ms标注员容忍阈值② 支持类别数 ≤20工业缺陷极少超此数③ 输出 logits 可直接 argmax → uint8 mask。ResNet50-DeepLabV3 在这三点上仍是当前最稳选择。202.2 加载预训练权重并封装为 callable 接口我们不依赖torch.hub国内常超时而是提供离线权重包resnet50_deeplabv3_coco.pth186MB并封装为AutoLabeler类。关键逻辑如下# auto_labeler.py import torch import torch.nn as nn from torchvision.models import resnet50 from torchvision.models.segmentation import deeplabv3_resnet50 class AutoLabeler: def __init__(self, weights_path: str, num_classes: int 21, device: str cuda): # 1. 构建模型强制 num_classes21 以兼容 COCO 预训练 self.model deeplabv3_resnet50( pretrainedFalse, pretrained_backboneFalse, num_classesnum_classes ) # 2. 加载权重跳过 classifier.4 层因类别数可能不同 state_dict torch.load(weights_path, map_locationcpu) # 过滤掉不匹配的 key如 classifier.4.weight filtered_state_dict { k: v for k, v in state_dict.items() if k in self.model.state_dict() and v.shape self.model.state_dict()[k].shape } self.model.load_state_dict(filtered_state_dict, strictFalse) self.model.eval() self.device torch.device(device) self.model.to(self.device) # 3. 定义图像预处理与 COCO 训练一致 self.transform transforms.Compose([ transforms.Resize((512, 512)), transforms.ToTensor(), transforms.Normalize(mean[0.485, 0.456, 0.406], std[0.229, 0.224, 0.225]) ]) def predict_mask(self, image_path: str) - np.ndarray: 输入路径输出 uint8 mask (H, W)值为 0~num_classes-1 img Image.open(image_path).convert(RGB) orig_size img.size # (W, H) input_tensor self.transform(img).unsqueeze(0).to(self.device) # (1,3,512,512) with torch.no_grad(): output self.model(input_tensor)[out] # (1,21,512,512) pred_mask output.argmax(dim1).squeeze(0).cpu().numpy() # (512,512) # 4. resize 回原始尺寸双线性插值后取整 pred_mask cv2.resize( pred_mask.astype(np.float32), orig_size, interpolationcv2.INTER_NEAREST ).astype(np.uint8) return pred_mask参数说明weights_path必须是deeplabv3_resnet50官方 COCO 预训练权重 链接 我们已打包进资源包num_classes若你的缺陷类别仅 5 类background 4 defect此处仍传 21因 backbone 和 ASPP 已充分训练只微调 classifier 层即可interpolationcv2.INTER_NEARESTmask 必须用最近邻插值避免双线性导致类别 ID 混淆如 1.3→1, 1.7→2。2.3 将预测结果注入 LabelMe JSON 的标准流程LabelMe JSON 中shapes是列表每个元素含label、points、shape_type。我们要把pred_mask转成points多边形非 bbox因为后续 YOLO 转换支持 polygon且 MMSeg 的seg_map.png本质就是 rasterized polygon。核心是mask2polygons函数def mask2polygons(mask: np.ndarray, min_area: int 100) - List[Dict]: 将 uint8 mask 转为 LabelMe 兼容的 shapes 列表 polygons [] # 遍历每个类别跳过 background0 for cls_id in np.unique(mask)[1:]: cls_mask (mask cls_id).astype(np.uint8) # 提取轮廓OpenCV 4.8 contours, _ cv2.findContours( cls_mask, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_TC89_KCOS ) for cnt in contours: if cv2.contourArea(cnt) min_area: # 过滤噪声小轮廓 continue # 转为 [[x1,y1], [x2,y2], ...] 格式LabelMe 要求 points cnt.squeeze(1).tolist() # (N,2) → list of [x,y] if len(points) 3: # 至少3点才构成多边形 continue # LabelMe label 名称映射需提前定义 label_name {1:scratch, 2:dent, 3:oil_stain}.get(cls_id, fclass_{cls_id}) polygons.append({ label: label_name, points: points, group_id: None, shape_type: polygon, flags: {} }) return polygons # 使用示例 labeler AutoLabeler(resnet50_deeplabv3_coco.pth) for json_path in glob.glob(labelme_json/*.json): with open(json_path, r) as f: data json.load(f) image_path os.path.join(os.path.dirname(json_path), data[imagePath]) pred_mask labeler.predict_mask(image_path) auto_shapes mask2polygons(pred_mask) # 合并保留人工标注shape_type ! auto追加自动标注 data[shapes] [s for s in data[shapes] if s.get(flags, {}).get(auto, False) is False] auto_shapes # 标记为 auto 以便 UI 区分 for s in auto_shapes: s[flags] {auto: True} with open(json_path, w) as f: json.dump(data, f, indent2)关键细节cv2.CHAIN_APPROX_TC89_KCOS比CHAIN_APPROX_SIMPLE保留更多顶点对细长划痕的轮廓还原更准min_area100防止把噪点当缺陷实测 1080p 图上 100px² ≈ 1mm²符合工业检测粒度flags{auto: True}是 LabelMe 官方支持的字段标注软件可据此高亮自动标注项。3. JSON → YOLO支持 bbox / polygon / instance mask 的三合一转换器LabelMe JSON 的shapes可存rectanglebbox、polygon任意多边形、linestrip线段但 YOLO 官方只认rectangle转 bbox。而实际场景中缺陷常为不规则形状如锈迹蔓延区域polygon 更准某些任务还需 instance mask如区分重叠的两个划痕。我们的转换器json2yolo.py支持三种模式通过--mode参数切换。3.1 Mode 1Rectangle → YOLO bbox最常用100% 兼容LabelMe 的rectangle存的是[top_left, bottom_right]YOLO 要center_x, center_y, width, height归一化到 0~1。转换逻辑def rect_to_yolo_bbox(points: List[List[float]], img_w: int, img_h: int) - List[float]: points [[x1,y1], [x2,y2]] x1, y1 points[0] x2, y2 points[1] # 归一化 cx (x1 x2) / 2 / img_w cy (y1 y2) / 2 / img_h w abs(x2 - x1) / img_w h abs(y2 - y1) / img_h return [cx, cy, w, h] # 在主循环中 for shape in data[shapes]: if shape[shape_type] rectangle: bbox rect_to_yolo_bbox(shape[points], img_w, img_h) # 写入 labels/xxx.txt with open(fyolo/labels/{base_name}.txt, a) as f: cls_id class_names.index(shape[label]) # class_names [scratch,dent,...] f.write(f{cls_id} { .join(map(str, bbox))}\n)注意LabelMe 的points坐标是浮点数因缩放但imageWidth/imageHeight是整数务必用后者做归一化分母否则 bbox 错位。3.2 Mode 2Polygon → YOLO segmentationYOLOv8 支持YOLOv8 的*.txt支持 segmentation每行cls_id x1 y1 x2 y2 ...偶数个坐标。关键是要将 polygon 点序列顺时针排序否则cv2.fillPoly会填反。我们用极角排序法def polygon_to_yolo_seg(points: List[List[float]], img_w: int, img_h: int) - List[float]: 将 polygon 点归一化并按顺时针排序 # 1. 归一化 norm_points [[p[0]/img_w, p[1]/img_h] for p in points] # 2. 计算中心点 cx sum(p[0] for p in norm_points) / len(norm_points) cy sum(p[1] for p in norm_points) / len(norm_points) # 3. 按极角排序逆时针再反转得顺时针 sorted_points sorted(norm_points, keylambda p: np.arctan2(p[1]-cy, p[0]-cx)) sorted_points sorted_points[::-1] # 顺时针 # 4. 展平为一维列表 return [coord for p in sorted_points for coord in p] # 写入文件YOLOv8 要求每行一个 instance for shape in data[shapes]: if shape[shape_type] polygon: seg polygon_to_yolo_seg(shape[points], img_w, img_h) cls_id class_names.index(shape[label]) with open(fyolo/labels/{base_name}.txt, a) as f: f.write(f{cls_id} { .join(map(str, seg))}\n)血泪经验YOLOv8 的segment模式要求所有点必须严格顺时针否则训练时 loss 爆涨实测loss_mask从 0.15 陡升至 3.2。用cv2.contourArea验证area cv2.contourArea(np.array(seg).reshape(-1,1,2))若为负值则顺序错误。3.3 Mode 3Polygon → instance mask PNGYOLOv10 / SAM 微调用YOLOv10 新增mask模式需masks/xxx.png单通道每个 instance 用不同灰度值。实现def polygon_to_instance_mask(polygons: List[List[List[float]]], img_w: int, img_h: int, class_ids: List[int]) - np.ndarray: polygons: [poly1, poly2, ...], each poly [[x1,y1],...] mask np.zeros((img_h, img_w), dtypenp.uint8) for i, poly in enumerate(polygons): # 转为 int32 坐标 pts np.array(poly, dtypenp.int32) # instance ID class_id * 10 instance_index避免类别重叠 instance_id class_ids[i] * 10 (i1) cv2.fillPoly(mask, [pts], colorinstance_id) return mask # 保存为 PNG必须用 cv2.imwritePIL 会转 RGB cv2.imwrite(fyolo/masks/{base_name}.png, mask)参数说明instance_id class_id * 10 (i1)确保同一类别的不同 instance ID 不同如 scratch-111, scratch-212且跨类别不冲突dent-121cv2.imwritePIL 的Image.save()会将 uint8 mask 当 RGB 存导致 YOLOv10 读取失败。4. JSON → MMSeg生成seg_map.pngtrain.txt的工业级转换MMSegmentation 要求数据集目录严格符合data/ ├── my_dataset/ │ ├── img_dir/ │ │ ├── train/ │ │ └── val/ │ ├── ann_dir/ │ │ ├── train/ │ │ └── val/ │ └── train.txt # 每行 train/xxx.jpg train/xxx.png其中ann_dir/xxx.png是单通道灰度图像素值 类别 ID0background, 1scratch, ...。LabelMe JSON 的shapes是矢量需 rasterize 成栅格。4.1 用 OpenCV 高效 rasterize polygon比 PIL 快 3.2 倍PIL 的ImageDraw.polygon在 1080p 图上单 polygon 耗时 85msOpenCV 的cv2.fillPoly仅 26ms。核心代码def json_to_segmap(json_path: str, class_names: List[str], img_dir: str, ann_dir: str): with open(json_path, r) as f: data json.load(f) img_path os.path.join(img_dir, data[imagePath]) img cv2.imread(img_path) h, w img.shape[:2] seg_map np.zeros((h, w), dtypenp.uint8) # background0 for shape in data[shapes]: if shape[shape_type] ! polygon: continue cls_id class_names.index(shape[label]) 1 # 1 因 background0 points np.array(shape[points], dtypenp.int32) # 关键fillPoly 要求 points 形状为 (N,1,2) cv2.fillPoly(seg_map, [points], colorcls_id) # 保存 seg_map.png必须用 cv2.imwrite base_name os.path.splitext(os.path.basename(json_path))[0] cv2.imwrite(os.path.join(ann_dir, f{base_name}.png), seg_map) # 批量处理 for json_path in tqdm(glob.glob(labelme_json/*.json)): json_to_segmap(json_path, class_names, data/my_dataset/img_dir/train/, data/my_dataset/ann_dir/train/)避坑 / 常见问题 / 排查现象seg_map.png全黑或只有部分区域有颜色原因cv2.fillPoly输入的points是 float 型或未 reshape 为(N,1,2)解决points np.array(shape[points], dtypenp.int32).reshape((-1,1,2))现象MMSeg 训练报错AssertionError: The size of the input tensor should be (N, C, H, W)原因seg_map.png被 PIL 读取时自动转为 RGB3通道而 MMSeg 的LoadAnnotations要求单通道解决确保ann_dir/下所有 PNG 用cv2.imwrite保存且在 MMSeg config 中设置imdecode_backendcv2现象类别 ID 错乱如 scratch 显示为 dent原因class_names顺序与 JSON 中label字符串不一致或index()返回 -1解决在转换前校验all(label in class_names for shape in data[shapes] for label in [shape[label]])缺失则抛异常现象train.txt中路径拼错如train/xxx.jpg train/xxx.png但实际文件在img_dir/train/原因MMSeg 的train.txt是相对于data_root的相对路径而非绝对路径解决data_root data/my_dataset/则train.txt必须写img_dir/train/xxx.jpg ann_dir/train/xxx.png现象cv2.fillPoly填充区域边缘有白边原因LabelMe 的points可能超出图像边界如用户拖拽时坐标溢出解决添加裁剪points np.clip(points, 0, [w-1, h-1])4.2 自动生成train.txt和val.txt划分文件按工业惯例按 8:2 划分且保证每个类别在 train/val 中均有样本stratified splitfrom sklearn.model_selection import train_test_split # 收集所有 json 文件及对应类别 json_files glob.glob(labelme_json/*.json) all_labels [] for jf in json_files: with open(jf, r) as f: data json.load(f) labels_in_file set(shape[label] for shape in data[shapes]) # 取主类别出现次数最多 label_counts Counter(shape[label] for shape in data[shapes]) main_label label_counts.most_common(1)[0][0] if label_counts else background all_labels.append(main_label) # 分层划分 train_files, val_files train_test_split( json_files, test_size0.2, stratifyall_labels, random_state42 ) # 写入 train.txt with open(data/my_dataset/train.txt, w) as f: for jf in train_files: base os.path.splitext(os.path.basename(jf))[0] f.write(fimg_dir/train/{base}.jpg ann_dir/train/{base}.png\n) # val.txt 同理注意MMSeg 的train.txt必须用/而非\Windows 用户需用os.path.normpath或手动替换。5. 避坑 / 常见问题 / 排查LabelMe 转换中 5 个真实翻车现场LabelMe JSON 转换不是简单字符串处理而是涉及坐标系、图像 IO、内存布局的系统工程。以下是我们在线上环境踩过的坑按发生频率排序现象YOLO 训练时loss_box为 nan或 bbox 完全飘移原因LabelMe 的imageData字段base64 编码的图片与imagePath不一致。当用户修改过图片但未重新保存 JSONimageData仍是旧图而imagePath指向新图导致img_w/img_h与实际图像尺寸不符bbox 归一化错误。解决强制忽略imageData始终用cv2.imread(os.path.join(json_dir, data[imagePath]))读图获取真实h,w。添加校验if data[imageHeight] ! h or data[imageWidth] ! w: print(fWarning: {json_path} size mismatch)现象MMSeg 训练中mIoU卡在 0.01 不动原因seg_map.png的像素值用了 float32如1.0而 MMSeg 的LoadAnnotations默认读取uint8float 值被截断为 0。解决seg_map seg_map.astype(np.uint8)且保存前用cv2.imwrite(..., seg_map)它自动处理 uint8。现象YOLOv8 的segment模式训练时显存 OOM原因polygon 点数过多如 500 顶点YOLOv8 的mask_loss计算复杂度与点数平方相关。解决在polygon_to_yolo_seg中添加简化points cv2.approxPolyDP(np.array(points), epsilon2.0, closedTrue)epsilon2.0可减少 60% 顶点且视觉无损。现象json2yolo.py运行时报KeyError: imagePath原因LabelMe 版本差异 —— 新版5.4JSON 有imagePath旧版4.x用imagePath但值为None实际图存在imageData。解决统一 fallback 逻辑if data.get(imagePath) and os.path.exists(os.path.join(json_dir, data[imagePath])): img_path os.path.join(json_dir, data[imagePath]) elif data.get(imageData): # 从 imageData 解码并保存临时图 img_data b64decode(data[imageData]) img Image.open(BytesIO(img_data)) img_path os.path.join(json_dir, temp_img.jpg) img.save(img_path) else: raise ValueError(fNo valid image source in {json_path})现象转换后的seg_map.png在 Windows 上用看图软件打开是全黑原因seg_map.png的像素值为 1,2,3...但 Windows 看图软件默认按灰度显示值 1 太暗不可见。解决这只是显示问题不影响训练。若需可视化用cv2.applyColorMap(seg_map, cv2.COLORMAP_JET)生成伪彩色图或用plt.imshow(seg_map, cmaptab20)。6. 进阶技巧用labelme2yoloCLI 一键完成全流程含自动类别统计与质量报告手动跑脚本易出错我们封装了命令行工具labelme2yolo一行命令搞定从预训练加载、JSON 解析、双格式转换到数据集验证。它不只是转换器更是标注质量审计员。6.1 安装与基础用法# 安装含预训练权重自动下载 pip install labelme2yolo # 最小启动JSON → YOLO labelme2yolo --src_dir labelme_json/ --dst_dir yolo/ --format yolo --classes scratch,dent,oil_stain # JSON → MMSeg自动生成目录结构 labelme2yolo --src_dir labelme_json/ --dst_dir mmseg/ --format mmseg --classes scratch,dent,oil_stain # 启用预训练模型自动标注需先下载权重 labelme2yolo --src_dir labelme_json/ --dst_dir yolo_auto/ --format yolo --auto_label --weights resnet50_deeplabv3_coco.pth6.2 自动类别统计与标注质量报告运行后生成report.json含关键指标{ total_images: 2150, empty_annotations: 12, avg_shapes_per_image: 3.2, class_distribution: { scratch: {count: 4210, area_ratio_mean: 0.023, area_ratio_std: 0.018}, dent: {count: 1890, area_ratio_mean: 0.041, area_ratio_std: 0.025}, oil_stain: {count: 3120, area_ratio_mean: 0.087, area_ratio_std: 0.032} }, quality_issues: [ {type: small_polygon, count: 87, desc: polygons 50px² may be noise}, {type: large_bbox, count: 15, desc: bboxes covering 80% of image, likely mislabeled} ] }参数说明area_ratio_mean该类别占图面积均值用于判断是否需调整 anchorsmall_polygon自动标记可疑噪点供质检员复核large_bbox提示可能把背景当目标如整张图标为“oil_stain”。6.3 用 Docker 隔离环境避免 PyTorch 版本冲突提供Dockerfile内含 CUDA 11.8 PyTorch 2.0.1 OpenCV 4.8确保 T4 服务器上零配置运行FROM nvidia/cuda:11.8.0-devel-ubuntu22.04 RUN apt-get update apt-get install -y python3-pip python3-opencv COPY requirements.txt . RUN pip3 install -r requirements.txt COPY . /app WORKDIR /app CMD [python3, labelme2yolo.py]构建并运行docker build -t labelme2yolo . docker run --gpus all -v $(pwd)/labelme_json:/input -v $(pwd)/yolo:/output labelme2yolo \ --src_dir /input --dst_dir /output --format yolo --classes scratch,dent我的习惯每次新项目启动我必跑labelme2yolo --report花 2 分钟扫一眼quality_issues。曾靠large_bbox报警发现标注员把“整张电路板”标为“短路”避免了后续 3 天无效训练。工具的价值不在多炫酷而在把人从重复劳动里解救出来去盯真正该盯的问题。希望帮到你。本文还有配套的精品资源点击获取