在笔记本上自制39M小语言模型指南
我的模型huggingface仓库地址:https://huggingface.co/Duoia/duogpt-40m-v1
我花了两天时间在我的笔记本上制作了一个会说英语,能回答简单问题和写儿童故事的小语言模型,名字叫DuoGPT,虽然没有什么用,但是写下来也许能帮助机器学习的小萌新(现在真的还有看园子的萌新吗),so,还请大佬轻喷
模型成本
这是DuoGPT的成本参数:
- RTX4060 Laptop (有8G显存,实际上用不到这么多)
- 预训练3.37小时(让它学会说英语)
- SFT2.07小时(让它会接话)
- 给梁文峰的大肥鱼投喂的五块钱
模型成果
DuoGPT最后的风格是这样的:
Q:/s A is same as B. | Is A diffrent from B?
A:No.
Q: /s Kate is a cat. | What is Kate?
A: Kate is a cat.
Q:/t Write a story about a shy cat.
A:Once upon a time,there was a shy cat named Kitty. Kitty was very shy and did not like to talk to other animals. One day, Kitty saw a little bird who wassad.The bird had lost its way and could not find its family. Kitty wanted to help the bird, so she decided to find its family.
Kitty and the bird walked and walked until they found the bird's family. The bird was so happy and thanked Kitty for helping. The bird's family was very grateful and gave Kitty a big hug. From that day on, Kitty was not so shy anymore. She made many new friends and had lots of fun. The moral of the story is that even if you are scared, you can still be brave and help others.
Q:/t Summarize the following story. I The bunny meets a small bird in the garden. The bird flies and sings. They say hello to each other. Next, the bunny sees a little bee flying around the flowers. It also finds a beautifulbutterfly in the sky. All the small animals are playing in the garden. After playing outside, the bunny feels hungry. It eats fresh grass and sweet carrots, then drinks some clean water. Soon,it meets a brown squirrel with a big tail."The squirrel can jump very high. The bunny plays and jumps with the squirrel happily.
A:A bunny makes friends with a small bird and a bee in the garden,playing and having fun in the garden.
总之这么小的模型能学会这么多东西已经是很强了,说实话我其实预计它连语法都学的颠三倒四,只能拼凑一些单词,能流畅表达意思还是太超模了
模型架构
模型的实际参数是38854144参数,是litgpt0.5.13框架,上下文是512,架构是11层深度,512宽度,8个头,head dim是64,SwiGLU inter是1368,用了RMSNorm,RoPE,没有bias,用了权重绑定
切词器大小是8192,其实这个切词器有点大了,对于这次的儿童故事,6k差不多就够了,切词器开太大一会导致切的比较碎,二会导致占太多权重,我们小模型这点权重是要精打细算的,另外我在设计特殊token时没有
训练语料
预训练语料的大头是TinyStoriesV2-GPT4-train,这是个由GPT4生成的儿童故事集,大概是270多万篇故事,儿童故事词汇量简单,语法简答,逻辑简单,适合我们的预训练,它的总大小是2.23G,然后我还混了五个切片的Children-Stories,大概是44.8万篇故事,它的词汇量比TSV2要稍微复杂一点,总大小是0.92G,这些数据集都能在huggingface下到,合计是7.45亿token(用我的切词器算的),接着整形,用litdata打成了513token的块
SFT语料70%来源是TinyStories自带的SFT,画风大概就是“写一个故事,以什么什么为主角”这类的,另外30%来源是bAbI,这是一个小任务数据集,大概就是给你一段上下文,然后问你一个问题,总共合计一万条
训练环节
接下来就是喜闻乐见的预训练环节,模型太小了,过度训练会过拟合,所以我只用了1 epoch,优化器挑的是AdamW,批大小是256,val曲线也是十分优雅的一直缓慢下降,最后loss停在了1.602,ppl停在了4.96,花了三个多小时,显存峰值是5.35G
SFT一开始只用了TinyStories的自带SFT,这个后果是无论说什么它都会写一段故事,所以后面又重新加训了bAbI(本来想用qwen2.53B小模型直接生成的,但是速度太慢水平还不高),ppl降到了三点多,说明了平常搞模型一定要随时备份checkpoint,不然预训练得重来了。。
总结
小模型一定要精打细算,比如分词就自己弄,不要用现成的,然后训练数据完全决定模型上限,务必根据自己的场景好好挑,下一步我准备加vikidia百科的语料,他们都是用很简单的英语写给小孩的,看看能不能让模型扩充一点知识
(我的模型huggingface仓库地址:https://huggingface.co/Duoia/duogpt-40m-v1)
浙公网安备 33010602011771号