并发编程(七):volatile——从语言规则到 CPU

目录


0. 这一篇继续回答什么?

前两篇从语言规则一路看到 CPU,解释了 Atomic 如何提供 Atomicity、Visibility 和 Ordering。

这一篇继续沿用同一条思路,只是对象换成 volatile。

volatile 不解决 counter++ 这样的复合更新。它更适合下面这种发布 / 观察关系:

Thread A                         Thread B

counter = 1                     if (ready) {
ready = true                        print(counter)
                                }

我们希望保证:

如果 Thread B 已经读到 ready == true,那么随后读取 counter 时必须看到 1。

这一篇先从 Java 的语言规则解释为什么成立,再继续往下看 HotSpot 和 CPU 如何实现。


1. Java volatile 在语言层保证什么?

1.1 Visibility 和 Ordering

把 ready 声明为 volatile:

int counter = 0;
volatile boolean ready = false;

// Thread A
counter = 1;
ready = true;

// Thread B
if (ready) {
    System.out.println(counter);
}

只有 ready 是 volatile,counter 仍然是普通变量。

JLS §17.4.4 规定:

“A write to a volatile variable v synchronizes-with all subsequent reads of v by any thread.”

因此,如果 Thread B 的 volatile read 观察到了 Thread A 写入的 ready = true,就有:

Thread A                              Thread B

counter = 1
    │
    │ program order
    ▼
volatile ready = true ── synchronizes-with ──► read volatile ready == true
                                                   │
                                                   │ program order
                                                   ▼
                                               read counter

再根据 happens-before 的传递性:

write counter
    ↓ happens-before
write volatile ready
    ↓ synchronizes-with
read volatile ready
    ↓ happens-before
read counter

所以当 B 已经读到 ready == true 时,前面的 counter = 1 对 B 可见,并且不能再出现:

ready == true
counter == 0

这里同时得到两个保证:

  • Visibility:A 在 volatile write 之前完成的写入,对随后观察到该 volatile write 的 B 可见;
  • Ordering:这些操作必须按照 happens-before 关系被观察,不能得到违反该顺序的结果。

1.2 volatile 不能保证复合操作的 Atomicity

例如:

volatile int counter = 0;

counter++;

counter++ 仍然是一个 Read-Modify-Write:

volatile read
      ↓
     add
      ↓
volatile write

volatile 可以约束这次读和写的可见性与顺序,但不能把整段 read → add → write 合成一次不可分割的更新。

多个线程同时执行时,仍然可能发生丢失更新。

如果需要原子的 RMW,应该使用 AtomicInteger、CAS 或 Mutex。


2. HotSpot 如何实现 volatile?

2.1 分层

和前面的 Mutex、Atomic 一样,先把实现层次固定下来:

Java Source Code
实例:volatile boolean ready
作用:声明具有 volatile 内存语义的字段
        │
        ▼
JVM Bytecode
实例:getfield / putfield(或 getstatic / putstatic)
字段元数据:ACC_VOLATILE
作用:表达字段访问;volatile 属性来自字段元数据
        │
        ▼
JVM Implementation(HotSpot)
实例:field->is_volatile() / MO_SEQ_CST / MemBar
作用:保留并实现 volatile 的内存顺序语义
        │
        ▼
x86-64 Hardware
实例:Load / Store / Cache Coherence / Fence
作用:提供 Visibility 与 Ordering 的硬件基础

这里和 synchronized、AtomicInteger 都不一样:

synchronized
  → 有 monitorenter / monitorexit 专用字节码

AtomicInteger
  → 普通 invokevirtual
  → 没有 Atomic 专用字节码

volatile
  → 仍然是 getfield / putfield
  → 没有 volatile load / store 专用字节码
  → volatile 属性记录在字段的 ACC_VOLATILE 元数据中

2.2 一次 volatile write / read 的完整实现路径

继续使用同一个 counter / ready 例子:

counter 的普通写入和 ready 的 volatile write / read 如何经过 JVM Bytecode、HotSpot,最终落实到 x86-64 Hardware。

主流程里最重要的不是某一条固定机器指令,而是 HotSpot 必须识别 ready 是 volatile 字段,并把它的内存语义一直保留到目标架构。

2.3 Runtime / Language Implementation 层的两种保证

2.3.1 Visibility

HotSpot 在解析字段访问时会区分普通字段和 volatile 字段。当前 C2 路径中,volatile 字段访问会带上 MO_SEQ_CST 一类的内存顺序信息,而普通字段使用普通的无序访问语义。

可以抽象成:

普通字段
  → ordinary load / store

volatile 字段
  → volatile load / store
  → 保留跨线程可观察的内存语义

也就是说,HotSpot 不能把 volatile 访问当作普通字段访问随意优化掉、合并或跨越必要的同步边界。

2.3.2 Ordering

HotSpot 还需要限制 volatile 前后的内存操作重排。

C2 的模型中可以看到典型的:

volatile store
  ← MemBarRelease 等顺序约束

volatile load
  → MemBarAcquire 等顺序约束

对于 volatile store → volatile load 这类还需要更强顺序的场景,还要建立相应的 StoreLoad / volatile barrier 约束。

这些是 HotSpot 的中间表示和编译约束;经过优化后,不代表每个节点都会一一对应成独立的 CPU fence。

相关实现可以查看 OpenJDK 的 parse3.cpp、memnode.hpp 和 templateTable_x86.cpp。

2.4 Hardware 层的两种保证

继续向下,最终还是回到第一篇的两类硬件能力:

Visibility
  → Cache Coherence
  → 让其他 CPU Core 不再长期使用已经失效的旧 Cache Line

Ordering
  → CPU Memory Ordering + Fence / 等价顺序约束
  → 保证 volatile 边界要求的内存访问顺序

在 x86-64 上,内存模型本身已经提供较强的顺序约束,所以很多 Acquire / Release 约束不需要额外生成一条硬件 fence。

真正需要 StoreLoad fence 时,HotSpot 的 Linux x86 路径可以使用:

lock; addl $0, 0(%rsp)

来建立 full fence 效果。

相关实现可以继续查看 OpenJDK 的 orderAccess_linux_x86.hpp。


3. Go 和 CPython 为什么没有 Java volatile?

Go 和 Python 都没有 Java 这种字段级 volatile 机制。原因不同,分别来看。

3.1 Go:这是语言设计选择

Effective Go 用一句话概括 Go 的并发设计:

“Do not communicate by sharing memory; instead, share memory by communicating.”

这句话反映的是 Go 的设计取向:并发同步应该通过显式的并发原语表达,而不是通过字段修饰符隐式表达。

因此,Go 没有像 Java 那样提供字段级 volatile。同类问题交给 Channel、Mutex 和 Atomic 处理。

对于和 Java volatile 最接近的场景,Go 使用 sync/atomic。The Go Memory Model 明确写道:

“This definition provides the same semantics as C++'s sequentially consistent atomics and Java's volatile variables.”

还是用前面的 counter / ready 例子:

var counter int
var ready atomic.Bool

// Goroutine A
counter = 1
ready.Store(true)

// Goroutine B
if ready.Load() {
    fmt.Println(counter)
}

如果 ready.Load() 观察到 ready.Store(true),那么两个 Atomic 操作之间建立 synchronized-before:

Goroutine A                          Goroutine B

counter = 1
    │
    │ sequenced-before
    ▼
ready.Store(true) ── synchronized-before ──► ready.Load() == true
                                                 │
                                                 │ sequenced-before
                                                 ▼
                                             read counter

因此,当 B 读到 ready == true 时,前面的 counter = 1 对 B 可见。

3.2 CPython:同步一直由 Runtime 承担

Python 没有 Java volatile,核心原因不在语法,而在同步职责一直放在 CPython Runtime。

主要有两点。

1. 传统 CPython 通过 GIL 覆盖了大量 Visibility / Ordering 问题

CPython C API 文档写道:

“only a thread that holds the GIL may operate on Python objects or invoke Python’s C API.”

传统 CPython 中,大量 Python 对象操作都经过 GIL 串行化。

GIL 的释放与重新获取也是线程同步边界,因此它在实际效果上覆盖了很多 volatile 用来解决的 Visibility 和 Ordering 问题。

但两者不是同一个东西:

Java volatile
  → 字段级语言语义

CPython GIL
  → Runtime 级同步机制

2. Free-threaded CPython 仍然选择 Runtime Lock + Atomic

GIL 可以关闭以后,CPython 也没有新增字段级 volatile。

PEP 703 的方案是:

“This PEP proposes using per-object locks...”

同时,Runtime 内部还会使用 Atomic 操作。

所以路线只是从:

GIL

变成:

Per-object Lock + Atomic

同步仍然留在 Runtime,没有上移成 Python 的字段级语言语义。

应用层需要发布 / 观察时,使用同步对象。Python 的 threading.Event 文档写道:

“one thread signals an event and other threads wait for it.”

还是套回前面的 counter / ready:

counter = 0
ready = threading.Event()

# Thread A
counter = 1
ready.set()

# Thread B
ready.wait()
print(counter)

这里 Event 承担了 ready 的同步职责:

Thread A                         Thread B

counter = 1
    │
    ▼
ready.set()  ───────────────►  ready.wait() returns
                                  │
                                  ▼
                              read counter

所以 Python 应用层使用 Event / Lock,底层的 Atomic 和 Memory Ordering 留在 CPython Runtime。


4. volatile、Atomic 和 Mutex 的关系

把前面几篇放在一起,三者解决的问题并不相同:

volatile Atomic Mutex
核心问题 发布 / 观察共享状态 一次共享状态操作 一段临界区
Atomicity 不保证复合 RMW 保证单次 Atomic 操作 保证临界区互斥
Visibility 是 是 是
Ordering 是 是 是
典型场景 初始化完成 / 配置加载 / 停止标记 / 状态发布 计数 / Add / CAS / Swap 多步逻辑 / 多字段整体更新

可以把关系简化为:

volatile
  → 让一次发布与后续观察建立 Visibility / Ordering

Atomic
  → 让一次共享状态操作本身不可分割
  → 同时提供相应的 Visibility / Ordering

Mutex
  → 用 Atomic 等底层能力竞争锁状态
  → 再保护一整段临界区

它们不是“强弱不同的同一种工具”,而是在解决不同粒度的问题。


5. 下一篇:从 Mutex 到读写锁

Mutex 同一时刻只允许一个执行单元进入临界区。

如果大量操作只是读取共享状态,让读者之间也互斥就没有必要。

下一篇继续看读写锁:如何让多个读者同时进入,同时仍然保证写入互斥。


本文首发于 ThinkerQAQ 的个人博客,由作者本人同步发布。原文可能持续修订,最新版本请以个人博客为准。

posted @ 2026-10-03 21:08  ThinkerQAQ  阅读(20)  评论(0)    收藏  举报