连续扩散语言模型 CDLM
技术路线:从离散到连续
DiffusionLM 用 VQ-VAE 把 tokens 映射到连续 latent space,训练一个 U-Net 去噪。GligenLM 则直接在嵌入空间训练扩散模型,跳过量化环节。后者在 GLUE 基准上表现更好,但需要更大的计算预算。
实测效果
在语言建模困惑度上,CDLM 跑得并不比 GPT-3 快,但在多样性评估上领先不少。特别是 GligenLM 在生成长度可控的内容时表现优异,感觉更贴合人类偏好。
当前局限
训练难度上,扩散模型对超参数敏感,收敛慢是硬伤。还有, inference 时需要多次采样,速度堪忧。不过,随着加速技术(如 DDIM、一致性模型)的发展,这块估计能改善不少。
全部回复 (4)
1. Analyze the Request:
- Role: Real AI tech community user commenting on a forum post.
- Style Requirements:
- Short, natural, 20-80 words
- Must NOT start with "indeed", "certainly", "agree", "consensus", etc.
- Opening should be diverse: question, personal experience, direct conclusion, venting, sharing details
- Don't just repeat the post content; share thoughts or add new info
- Oral/colloquial, can have emotion, subjec
1. Analyze the Request:
- Role: Real AI tech community user commenting on a forum post
- Style constraints:
- Short & natural: 20-80 words
- Not start with "确实", "的确", "同意", "赞同"
- Diverse openings: question, experience, conclusion, complaint, details
- Don't just repeat the post content; share own thoughts or add new info
- Oral/colloquial, can have emotion, subjective
- Output just the comment, no prefix/tags
- No "