实验CPU密集型
1 如果vmstat r只有自己,那一个while是否能不让出CPU时间片,即没有非自愿切换?
import java.util.concurrent.CountDownLatch;
public class TestIO {
public static void main(String []f) throws Exception {
final CountDownLatch countDownLatch = new CountDownLatch(1);
Runnable runnable = new Runnable() {
@Override
public void run() {
try {
countDownLatch.await();
long st = System.currentTimeMillis();
for(long k=0;k<100000000000L;++k ) {
int i= 0;
int j= i*4;
}
System.out.println(System.currentTimeMillis() - st);
Thread.sleep(100000000);
} catch (InterruptedException e) {
e.printStackTrace();
}
}
};
for(int i=0; i<Integer.parseInt(f[0]); ++i) {
Thread thread = new Thread(runnable, "test" + i);
thread.start();
}
countDownLatch.countDown();
}
}
这个程序没有共享内存,没有IO

33秒16次被动切换,看上去没多少切换,所以时间片到时间了可能操作系统视情况不切换,或者干脆直接给予大时间片,因为时间片本来就可以是操作系统动态决定的
另一个问题,切换用的是内核时间片?
内核具有高level的内存访问权限,拷贝的工作必须由内核来做
2 us是不是恒定的?
2.1 物理机
4个CPU,18核,每个核心2线程,共144
| runtime | us | sy /100 | cs | ucs | |
| 1 | 1.38 | 1.36 | 0 | 1 | 2 |
| 10 | 1.40 | 1.38 | 0 | 77 | 1 |
| 50 | 1.84-23.5 | 2.33 | 0 | 70 | 3 |
| 100 | 4.4 | 4.03 | 2 | 313 | 18 |
| 143 | 3.9-5.0 | 3.99 | 3 | 273 | 191 |
| 144 | 6.1-7.3 | 5.58 | 5 | 292 | 209 |
| 4.4-5.6 | 4.98 | 2 | 221 | 169 | |
| 150 | 5-5.9 | 5.06 | 5 | 329 | 242 |
| 200 | 9.4 | 5-6 | 6 | 179 | 446 |
| 300 | 11.5 | 5 | 4 | 167 | 406 |
| 600 | 21.6 | 3.5 | 2 | 182 | 294 |
| 19.5 | 4-6 |
从单线程到CPU线程数,程序user time可能差4倍,即使没有共享内存
1thread

50threads

143threads

144


150threads

200


2.2 阿里云
2.1双核处理器
| runtime | us | sy | cs | ucs | |
| 1 | 33 | 33 | 0 | 4 | 16 |
| 2 | 67 | 67 | 0 | 4 | 1500 |
| 4 | 133 | 66 | 0 | 7 | 8500 |
| 8 | 273 | 68 | 0 | 12 | 6800 |
2threads

4threads

2.2四核
| runtime | us | sy | cs | ucs | |
| 1 | 33 | 33 | 0 | 5 | 18 |
| 2 | 66 | 66 | 0 | 5 | 9/58 |
| 3 | 68*1,100*2 | 68*1*100*2 | 0 | 6/7/8 | 20/64/114 |
| 4 | 132*4 | 130*4 | 0 | 6/7/8 | 12000 |
| 5 |
170*5 152-165 |
128*5 | 7 | 22000 | |
| 8 | 263*8 | 130*8 | 7 | 15000 | |
| 12 | 260-400 | 130*12 |
cpu用户态时间在线程数小于核数时不是恒定的
大于等于核数就恒定了;大于核数后开始出现cpu pending
等于核数开始,非自愿上下文切换显著增加,cpu开始调度
同样的代码,开4个线程,4核处理器cpu用户态花130秒,2核处理器只花65秒
用户态增加与ucs好像无关,比如4核1 2 线程切换数据几乎一样
从单线程到CPU线程数,差4倍,与1物理机结论一致

2threads
test1

test2

3threads
test1

test2

4threads
test1

test2

5thread

8

12

3 可能的原因
=================
https://cloud.tencent.com/developer/ask/sof/108558556
我对CPU时间的理解是,在同一台机器上,每次执行都应该是相同的。每次都需要相同数量的cpu周期。
但是我现在正在运行一些测试,执行一个基本的回声"Hello“,它给了我0.003到0.005秒的时间。
我对CPU时间的理解是错误的,还是在我的测量中出现了问题?
你的理解是完全错误的。在现代CPU上运行现代OSes的现实世界的计算机并不是简单的理论抽象.有各种各样的因素会影响CPU时间代码执行所需的时间。
考虑内存带宽。在一台典型的现代机器上,运行在机器核心上的所有任务都在争夺对系统内存的访问。如果代码同时运行在另一个核心上的代码正在使用大量的内存带宽,这可能导致访问RAM占用更多的时钟周期。
许多其他资源也是共享的,例如缓存。假设代码经常被中断,以便让其他代码在核心上运行。这将意味着,代码将经常发现缓存冷,并采取了大量的缓存错过。这也会导致代码占用更多的时钟周期。
让我们也来谈谈页面错误。代码本身可能在内存中,也可能在代码开始运行时不在内存中。即使代码在内存中,您也可能会或不可能使用软页错误(以更新操作系统对正在积极使用的内存的跟踪),这取决于该页上次发生软页错误的时间或加载到RAM中的时间。
您的基本hello world程序是对终端执行I/O操作。所需时间取决于当时与终端交互的其他内容。
https://cloud.tencent.com/developer/ask/sof/115064332
没用
https://cloud.tencent.com/developer/ask/sof/108487976
我已经创建了一个程序,该程序从argv获取参数,并为每个线程创建一个线程,线程关联设置为参数的int值。例如,./main 3 4将创建两个线程,第一个线程将在第三个cpu上运行,第二个线程将使用第四个cpu。
一个线程需要1秒才能完成(在数组int10000上执行数学操作),当我运行time ./main 1 2时,我看到了预期的1秒实时,但是当我运行time ./main 1 3时,我看到了2秒而不是1秒,我认为这与numa节点有关,但是time ./main 1 4导致了1秒的实时。
经过更多的测试后,我发现只有1对3对和2对4对花的时间是预期的两倍。而且,用户的时间也是原来的两倍。
$ time ./main 1 2 real 0m1.058s user 0m2.100s $ time ./main 1 3 real 0m2.019s user 0m4.016s $ time ./main 1 4 real 0m1.090s user 0m2.152s $ time ./main 2 4 real 0m2.014s user 0m4.016s $ time ./main 2 3 real 0m1.094s user 0m2.156s $ time ./main 3 4 real 0m1.170s user 0m2.316s
void math_ops() {
size_t len = 14800; // with this number it takes around 1s to compute on my hardware
int* abc = new int[len+1];
memset(abc, 7, len);
for(int i = 1; i < len; i++) {
for(int j = 1; j < len; j++) {
abc[i] *= abc[j];
abc[j+1] -= abc[i-1];
abc[j-1] -= abc[i+1];
}
}
}
int main(int argc, char** argv) {
std::vector<std::thread> vec(argc);
int thread_num = argc - 1;
for (int i = 0; i < thread_num; i++) {
std::thread t(math_ops);
// sets thread affinity equal to the second parameter
set_affinity(t, atoi(argv[i+1]) - 1);
vec[i] = std::move(t);
}
for (int i = 0; i < thread_num; i++) {
vec[i].join();
}
return 0;
}
这可能是由于超线程。你看到的四个核心并不是真正的4个核心,它们可能只是两个核心,有两倍多的执行单元。这意味着在属于同一物理核心的虚拟核上运行的两个线程必须共享该核心的一些资源。
当您在两个不同的物理核上运行时,没有资源共享,代码执行得更快。
您可以通过阅读/sys/devices/system/cpu/cpu0/topology/thread_siblings_list (将cpu0替换为任何其他core#)来找出哪些核心是兄弟姐妹。
浙公网安备 33010602011771号