翻译[10]-使用Python构建C编译器
使用Python构建C编译器
- 原文地址: [https://www.reddit.com/r/Python/comments/eieuld/c_compiler_written_in_python/]
- 原作者: Reddit 用户 u/The_Acronym_Scribe,发表于 r/Python,2019 年 11 月(12 分 / 10 条评论,帖子已锁定归档)
- 存档时间: 2026-10-05(SingleFile 保存)
译者注:这张帖子是本博主后续折腾 Python 编译 C 的线索来源——评论区里
pfalcon2 推荐的 PPCI("Python 版的 LLVM")与 PeridexisErrant 推荐的
CSmith(随机 C 程序生成器,John Regehr 出品),后来都被下载进了
ttf_window工程(快照ppci-master/与csmith-master.zip)实际试玩,
详见同系列《折腾笔记[64]-使用python编译c语言代码解析ttf字体》。原帖中的
仓库如今也有变迁:CarterTS/CCompiler已改名为
AshTS/CCompiler(最后一次提交停在
2020-01-01,作者"周更/双周更"的计划并未持续);ppci-mirror已更名为
windelbouwman/ppci,截至 2025 年
10 月仍在维护。
用 Python 写了个 C 编译器
50 天前,当我发现这个十年只剩 50 天时,我决定终于要动手做一个已经计划了好
几个月的项目。很高兴地告诉大家,我完成了它的第一阶段:它已经能编译 C 语言
的大部分内容——https://github.com/CarterTS/CCompiler(链接现重定向至
AshTS/CCompiler)。
不过有几点需要说明:这并不是一个完整的 C 编译器,还不支持浮点数、联合体
(union)、枚举(enum)、位域(bit field),另外还有若干较小的不一致之处。
此外,这个编译器不会输出可供链接的 ELF 机器码文件,而是输出汇编。默认配置
下,它为我专门为本项目做的一个模拟器生成汇编,这个模拟器同时充当测试用例:
https://github.com/CarterTS/DemoProcessor。
这正是我想向各位求助的地方:这个编译器远非完美,而我发现最有效的测试方法
就是不断往它上面砸代码,看哪里会崩,然后修掉这些 bug。如果能收到关于代码、
项目整体、尤其是功能方面的反馈,我会非常感激;往这两个仓库提 issue 更是
帮了大忙。
假期即将结束,再过几天就要返校,这个项目的进度会有所放缓。不过我会尽量
保持每周或每两周一次更新,并多少写一点开发日志。
欢迎大家提问,当然也万分感谢任何反馈。祝大家新年快乐,愿你们的代码没有
bug!
评论(按"最佳"排序)
oar335 · 7 分
我自己也写过几个玩具编译器,其中也包括一个用 Python 写的、面向 x86 的
C 编译器(准确说是 C 的子集)。事后回想,我当初应该直接面向 LLVM 的。
我注意到你的解析器是手写的。以我的经验,那是整个代码里最脆弱的部分,写起来
也很枯燥。推荐使用解析器生成器(比如 flex/bison 或 ANTLR)。
flex/bison 的 Python 绑定见:http://freenet.mcnabhosting.com/python/pybison/
译者注:该个人站点现已无法访问。今天的 Python 生态里可以改用
ANTLR4 的 Python 运行时或纯 Python 的
SLY / PLY。
[已删除用户] · 4 分
Guido van Rossum 也推荐用解析器生成器。
The_Acronym_Scribe(原发帖人) · 1 分
我完全不知道 flex 和 bison 还有 Python 实现,太感谢了。
另外,我确实应该面向 LLVM。等我把它做成一个真正的编译器(输出 ELF 文件)
时,肯定会走这条路。
dbramucci · 2 分
我还没细看你的项目,但如果你想再多测测它,可以考虑基于属性的测试
(property based testing)。当然,为编译器设计"属性"挺有挑战的,以下是我
想到的一些技巧。(参考:Zac Hatfield-Dodds 关于用 Hypothesis 做属性测试的
演讲,以及 Hypothesis 属性测试库)
- 预言机属性(Oracle Properties)——可以设想这样一条性质:"如果一个
程序能被 CCompiler 编译,那它也应该能被 gcc 编译"(也就是说,如果它放过了
gcc 不接受的程序,那多半是 bug)。(难点在于测试框架可能很难随机生成足够
多可编译的字符串,导致这条性质效果有限。)同理,重构时你可以测试"上一版
实现"与"当前版本"行为完全一致;加优化时则测试"优化前后代码行为完全相同"
(注意:如果你利用未定义行为(UB)做激进优化,这条就不一定成立了)。 - 蜕变属性(Metamorphic Properties)——直接写性质很难,但我们可以对
一个能编译/运行的程序做一些结果可预测的改动,例如:-
如果程序 A 能编译,那么把任意符号 x 重命名为符号 y 应当:
- 不改变可执行文件的输出;
- 只把汇编里的 x 换成 y,其余一概不变。
-
在任何已有空白处增加空白字符,汇编不应改变(单行注释除外:
//hello world和//hello换行world含义不同)。 -
更一般化的例子:拿一个手写的程序,把所有与输出无关的部分消去,确认输出
按预期变化。(下面用<x>表示把变量 x 的文本插入此处;我本想像 f-string
那样用{x},但那样就得转义 C 代码里的花括号了。)int fib(int n) { if (n <= 1) { return 1; } else { return fib(n-1) + fib(n - 2); } }对任意满足 x != y、字母数字混合且不以数字开头的字符串 x、y,它可以变成:
int <x>(int <y>) { if (<y> <= 1) { return 1; } else { return <x>(<y>-1) + <x>(<y> - 2); } }输出中唯一的变化应当是:符号
fib的每一次出现都变成你随机生成的 x。
如果你的编译器做不到这一点,那它几乎肯定有 bug。
-
- 反向预言机(Reverse Oracle)——写一个把你的 AST 翻译回 C 程序的
"编译器"(简单的递归文本替换就够了)。这样测预言机属性会容易得多(你还
可以在"AST→C"之后对生成代码做随机文本替换,以获得很好的分布)。- 解析"AST→C"产物后得到的 AST,应当与你最初的 AST 相同(或等价)。
总之,我列出这些,是因为上次我给编译器写属性测试时,很希望能有这样一份
清单。以我的经验,属性测试非常擅长发现两类问题:要么是你对问题本身理解有误
("哦,不是任意字母数字串,而是字母开头的字母数字串"),要么是真正的 bug
("糟糕,如果 C 文件没有函数、只有常量,我的编译器就崩了")。希望这些对你
有帮助。
另外,如果你想找简单的 C 程序(版权问题需要你自己核实),可以看看
Benchmarks Game、
Codewars 和 Project Euler
这类刷题/游戏网站,也可以针对你已实现的特性找 C 教程和 Stack Overflow 上的
代码。遗憾的是我没有逐一核实这些网站对用户提交代码的版权政策,所以这些样例
未必能直接加进你项目的测试集。
译者注:Benchmarks Game 的原链接现已 404,站点迁移到
https://benchmarksgame-team.pages.debian.net/benchmarksgame/,源码仓库在
https://salsa.debian.org/benchmarksgame-team/benchmarksgame/。
PeridexisErrant · 3 分
测 C 编译器?你应该用 CSmith!
John Regehr 在"自动挖掘编译器 bug 的工具"(主要针对 C)方面做了大量工作和
研究。把这些工具拿来,测这条性质:"任取一个 CSmith 生成的程序,我的编译器
产出的可执行文件,与(生产级 C 编译器的)行为等价"。强得简直像开挂 😃
The_Acronym_Scribe(原发帖人) · 1 分
哇,谢谢,我一定会研究一下。
The_Acronym_Scribe(原发帖人) · 2 分
非常感谢这份清单,我肯定会用上这些测试方法。
pfalcon2 · 1 分
这真的很棒!
不过这已经是我看到的第三、第四个用 Python 写的 C 编译器了,没有一个是完整
的,甚至没有一个能持续开发下去。
不如加入我们的 PPCI 项目吧,那才是真正先进的编译器项目——简直就是 Python
版的 LLVM:https://github.com/windelbouwman/ppci-mirror/(该仓库现已更名为
windelbouwman/ppci)。
GuybrushThreepwo0d
总比用 Python 写 Python 解释器强。
nathanjell
为什么不呢?这是个相当好的教学项目。
C Compiler Written in Python(原文存档)
Fifty days ago, upon discovering that there were 50 days left in the decade, I decided to finally get working on a project I have meant to work on for several months. I am happy to say that I have completed the first leg of something which can compile most of the C Programing Language: https://github.com/CarterTS/CCompiler.
However, there are several caveats here, this is not a complete C compiler, it is missing support for floats, unions, enumerations, bit fields, along with several other smaller inconsistencies. Furthermore, the compiler does not output machine code in ELF files for linking. Instead, it outputs assembly. By default, it is configured to output assembly for an emulator I built specifically for this project to act as a test case (https://github.com/CarterTS/DemoProcessor).
This is where I would like to reach out to you for assistance, this compiler is far from perfect, and the most productive method for testing it I have found is by throwing code at it, seeing what breaks and fixing those bugs. I would greatly appreciate feedback on the code, the project in general, and especially the functionality. I would be very grateful for issues filed on either of these repositories.
Since the holiday season is coming to a close and with the return to school approaching in a few days, work on this particular project will slow somewhat, however, I am going to attempt to get weekly or bi-monthly updates together, along with something of a dev blog.
I would love to answer any questions you all may have, and of course, feedback is greatly appreciated. I hope you all have a wonderful new year and your code may be bug-free!
Comments (sorted by "Best")
oar335 · 7 points
I've written a couple of toy compilers myself. I also wrote a C Compiler written in Python, targeting x86 (actually a subset of C). In retrospect I should have targeted llvm.
I noticed you hand-rolled your own parser. In my experience that was the most fragile part of the code and felt like a menial task to implement. I recommend using parser generators (i.e. flex/bison or ANTLR).
For python bindings for flex/bison see: http://freenet.mcnabhosting.com/python/pybison/.
[deleted] · 4 points
Guido Van Rossum also recommends using a parser generator
The_Acronym_Scribe (OP) · 1 point
I had no clue that there was a python implementation for flex and bison, thank you so much.
In addition, I really should have targeted LLVM, that is definitely the route I will take for when I work on making this a true compiler (to ELF files).
dbramucci · 2 points
Not that I have looked at your project but if you want to test it some more, consider property based testing. Of course, it is challenging to come up with properties for a compiler but here are a few techniques that come to mind. (A talk by Zac Hatfield-Dodds about property based testing with Hypothesis, and the Hypothesis Property based testing library)
- Oracle Properties — Consider a property like "If this program compiles with CCompiler, it should also compile with gcc" (that is if it allows a program to compile that gcc wouldn't it's probably a bug). (It may be hard for the testing framework to find a sufficient number of random strings that are compile-able for this to be effective). Likewise you can test that "last implementation" has exactly the same behavior as "current" if you are refactoring and "code has exact same behavior" as current if you are adding optimizations (note that this won't hold true if you are using UB to add optimizations aggressively).
- Metamorphic Properties — It is hard to come up with properties but we can think of changes to a compiling / runnable program that we can predict the resulting change for, examples include:
-
If Program A compiles then renaming any symbol x to a symbol y should
- Not change the output of the executable
- Change x in the asm to y but nothing else.
-
Adding any whitespace to any existing whitespace should not change the asm (with the exception of one line comments where
//hello worldmeans something different than//hello\nworld). -
Generalized example: take a program that you write by hand and eliminate all parts that are irrelevant to the output and ensure the output changes as expected. (I'll write
<x>to indicate a place to plug in the text for variablex(I would use{x}like in an f-string but I would then need to escape brackets in c code.)int fib(int n) { if (n <= 1) { return 1; } else { return fib(n-1) + fib(n - 2); } }could become for any alphanumeric (without leading number) strings x and y where x != y
int <x>(int <y>) { if (<y> <= 1) { return 1; } else { return <x>(<y>-1) + <x>(<y> - 2); } }the only change to the output should be every instance of the symbol
fibshould now become whatever your randomxwas. If your compiler doesn't do this, it's almost certainly bugged.
-
- Reverse Oracle — Write a "compiler" that translates your AST to a C program (just simple recursive text substitution should suffice). This can help you test Oracle Properties more easily (and you can do random text substitution to the resulting code after the "AST->C" transformation to get a really good distribution).
- The AST you get after parsing the "AST->C" output should be the same (or equivalent) to what you started with.
Anyways, I am just listing these because I would have liked a list like this the last time I tried writing property tests for a compiler. and in my experience Property tests can be very good at finding either things you misunderstood about your problem "oh, not alphanumeric strings but alphanumeric strings with a leading alpha character" or bugs "oops my compiler crashes if a c file has no functions and only has constants". Therefore, you can use this for some help.
Also, if you are looking for simple C-Programs (You'll have to check copyright issues yourself) you can check the benchmark games, game websites like codewars and Project Euler and C tutorials to features you have implemented and help websites like Stack Overflow. Unfortunately, I haven't checked each websites copyright policies on submitted code so the linked websites might not allow you to add these samples to your project as test cases.
PeridexisErrant · 3 points
Testing a C Compiler? You should use CSmith!
John Regehr has done a lot of work and research on tools to automatically find bugs in compilers, usually targeting C. Grab those, and test the property "Given any program from CSmith, by compiler produces an executable with equivalent behaviour to (production C compiler)"
It's practically cheating it's so powerful 😃
The_Acronym_Scribe (OP) · 1 point
Wow, thanks, I will definitely look into this.
The_Acronym_Scribe (OP) · 2 points
Thank you so much for the list, I am definitely going to make use of these testing methods.
pfalcon2 · 1 point
This is absolutely great!
But this is also third or forth C compiler written in Python I see, none complete or even actively developed past some time.
Come instead join us in the PPCI project, which is a truly advanced compiler project, literally LLVM in Python: https://github.com/windelbouwman/ppci-mirror/
GuybrushThreepwo0d
Better than writing a python interpreter in python
nathanjell
Why not? A pretty good educational project

浙公网安备 33010602011771号