OS-lab4
OS-lab4
本节重点在于理解系统调用与进程创建,进程通信。
系统调用
在这里我们首先需要区用户态和内核态等几组概念:
- 用户态和内核态(也称用户模式和内核模式):它们是
CPU运行的两种状态。根据Lab3的说明,在MOS操作系统实验使用的仿真R3000 CPU中,该状态由CP0 SR寄存器中KUc位的值标志。 - 用户空间和内核空间:它们是虚拟内存(进程的地址空间)中的两部分区域。根据
Lab2的说明,MOS中的用户空间包括kuseg,而内核空间主要包括kseg0和kseg1。每个进程的用户空间通常通过页表映射到不同的物理页,而内核空间则直接映射到固定的物理页以及外部硬件设备。CPU在内核态下可以访问任何内存区域,对物理内存等硬件设备有完 整的控制权,而在用户态下则只能访问用户空间。 - (用户)进程和内核:进程是资源分配与调度的基本单位,拥有独立的地址空间,而内核负责管理系统资源和调度进程,使进程能够并发运行。与前两对概念不同,进程和内核并不是对立的存在,可以认为内核是存在于所有进程地址空间中的一段代码。
为什么要这么区分呢?首先来说,对于操作系统的一系列设置肯定要放在内核态,因为其涉及到一些权限比较高的操作。另外,我们考虑用户态,也就是暴露给用户的接口肯定不能很底层,不然什么操作都允许的话容易乱搞把系统搞崩溃。因此产生了这两种状态。
但是有时候一些用户又不得不使用一些权限高的操作,那怎么办呢?可以事先在内核态中将其封装好,然后用户态直接 "调用" 封装好的接口即可。
系统调用机制的实现
这个调用究竟是如何实现的?实际上是通过中断异常处理来实现。
首先在用户态运行到某一个接口时,有这个接口出发异常,然后由该异常陷入内核态,跳转到相应的处理程序后处理,全部处理异常之后再返回到用户态即可。
对应到代码中,事实上内核态和用户态的代码分别放置于不同的文件夹下,在 /user 文件夹下放置的是用户态的代码,而在 /kern 中放置的是内核态的代码。两者在很多权限上不一致。例如:在标记当前进程信息的 curenv 就只可以在内核态访问不对用户态开放。同时调试输出的时候用法也不同。在用户态中输出使用的是位于 user/lib/debugf.c 中的函数,而内核态使用的则是之前写好的 printk 函数。下面我们对某一个系统调用跟踪一下。
以 user/lib/ipc.c 中的某一函数中的 syscall_yield() 为例,使用于:
void ipc_send(u_int whom, u_int val, const void *srcva, u_int perm) {
int r;
while ((r = syscall_ipc_try_send(whom, val, srcva, perm)) == -E_IPC_NOT_RECV) {
syscall_yield();
}
user_assert(r == 0);
}
该函数的功能是实现进程间的 ipc 通信机制。首先调用 syscall_ipc_try_send 检查需要通信的一方是否准备好,如果没有准备好则采用 轮询 的方式等待并切换进程,其中的 syscall_yield() 则为实现进程之间的切换,需要进入内核态实现。
其声明位于 user/include/lib.h 中:
void syscall_yield(void);
实现位于 user/lib/syscall_lib.c 中:
void syscall_yield(void) {
msyscall(SYS_yield);
}
int syscall_env_destroy(u_int envid) {
return msyscall(SYS_env_destroy, envid);
}
int syscall_mem_alloc(u_int envid, void *va, u_int perm) {
return msyscall(SYS_mem_alloc, envid, va, perm);
}
通过 msyscall 实现异常中断,其中第一个参数表示是第几号中断,后续的参数表示向该内核态中的处理程序传入的参数例如后面两个系统调用实现,而中断对应了内核态中相应的处理程序,其中的对应实现于内核态中的 ./include/syscall.h 中的一个枚举类型数组:
enum {
...
SYS_yield,
SYS_env_destroy,
...
SYS_ipc_try_send,
` ...
SYS_read_dev,
MAX_SYSNO,
};
最后实现位于 ./kern/syscall_all.c 中,首先根据数组找到对应的处理程序:
void *syscall_table[MAX_SYSNO] = {
...
[SYS_yield] = sys_yield,
[SYS_env_destroy] = sys_env_destroy,
...
[SYS_ipc_try_send] = sys_ipc_try_send,
...
[SYS_write_dev] = sys_write_dev,
[SYS_read_dev] = sys_read_dev,
};
然后进入到该文件中对应的处理程序中处理即可:
exercese 4.7
void __attribute__((noreturn)) sys_yield(void) {
// Hint: Just use 'schedule' with 'yield' set.
/* Exercise 4.7: Your code here. */
schedule(1);
}
当然,还有一个非常精妙的地方需要提及,请大家观看这个系统调用在用户态和内核态的参数传递情况:
// 用户态
int syscall_mem_alloc(u_int envid, void *va, u_int perm) {
return msyscall(SYS_mem_alloc, envid, va, perm);
}
// 内核态
int sys_mem_alloc(u_int envid, u_int va, u_int perm) {
struct Env *env;
struct Page *pp;
...
return page_insert(env->env_pgdir, env->env_asid, pp, va, perm);
}
在参数传递的时候,由用户态传递的虚拟地址 va 的参数类型本来是 void* 指针,但是进入内核态以后却编程了 u_int 类型,为什么要这样去设计呢?我们要根据两边的需求来分析。
用户程序里访问地址中的内容(比如往某一个数组里填东西),那都应该是 void* 的。而当用户程序希望分析地址时,一般就会转化为 u_int,比方说 fork 时的 duppage 以及 cow_entry 中也有 u_int 作为地址类型。在这个函数中,更多的是强调对这个虚拟地址调用创建一些事,并不是分析,因此使用了 void*,而内核态中真正实现起来则强调了 "分析地址"的功能:解析地址,为该地址做映射。
在C 语言里面,类型其实并不是一个很强的属性,所以我们说 C 是弱类型的语言。类型更多的是提供“数据的解释”功能,他并不限定某一段数据必须由某一特定类型来解释。因此强制类型转化、传参时类型变更,都是司空见惯的。比方说你写 qsort,你的 int [] 类型的数组,能隐式地转化为 void * 类型。这就是一种弱类型、类型变更的体现。所以阅读 C 代码时,不能对参数类型抓得过紧,应当要结合实际的意义来理解函数的功能。
那么接下来我们就具体对系统调用的实现进行分析。异常分发的具体流程如下:

当 user/lib/syscall_lib.c 中定义的用户包装函数 syscall_* 调用 msyscall 时,根据 以上的调用规范,需要传递给内核的系统调用号以及其他参数已经被合理安置。接下来我们只需要编写用户空间中的 msyscall 函数,这个叶函数没有局部变量,也就是说这个函数不需要分配栈帧,只需执行自陷指令 syscall 来陷入内核态并在处理结束后正常返回即可。
exercise 4.1
LEAF(msyscall)
// Just use 'syscall' instruction and return.
/* Exercise 4.1: Your code here. */
syscall
jr ra
END(msyscall)
在通过 syscall 指令陷入内核态后,处理器将 PC 寄存器指向一个内核中固定的异常处 理入口。在异常向量表中,系统调用这一异常类型的处理入口为 handle_sys 函数,它是在 kern/genex.S 中定义的对 do_syscall 函数的包装,我们需要首先在 kern/syscall_all.c 中 实现 do_syscall 函数。
exercise 4.2
void do_syscall(struct Trapframe *tf) {
int (*func)(u_int, u_int, u_int, u_int, u_int);
// 相当于是函数指针,对应了使用哪个函数处理异常
int sysno = tf->regs[4];
if (sysno < 0 || sysno >= MAX_SYSNO) {
tf->regs[2] = -E_NO_SYS;
return;
}
/* Step 1: Add the EPC in 'tf' by a word (size of an instruction). */
/* Exercise 4.2: Your code here. (1/4) */
tf->cp0_epc += 4;
/* Step 2: Use 'sysno' to get 'func' from 'syscall_table'. */
/* Exercise 4.2: Your code here. (2/4) */
func = syscall_table[sysno];
/* Step 3: First 3 args are stored in $a1, $a2, $a3. */
u_int arg1 = tf->regs[5];
u_int arg2 = tf->regs[6];
u_int arg3 = tf->regs[7];
/* Step 4: Last 2 args are stored in stack at [$sp + 16 bytes], [$sp + 20 bytes]. */
u_int arg4, arg5;
/* Exercise 4.2: Your code here. (3/4) */
arg4 = *(int*)(tf->regs[29] + 16);
arg5 = *(int*)(tf->regs[29] + 20);
/* Step 5: Invoke 'func' with retrieved arguments and store its return value to $v0 in 'tf'.
*/
/* Exercise 4.2: Your code here. (4/4) */
tf->regs[2] = (*func)(arg1, arg2, arg3, arg4, arg5);
}
其中我怎么知道哪个寄存器对应哪个位置呢?看图:

正好对应了结构体 Trapframe 中的数组 resg[32] 的每一位。这时候就要问了,既然是陷入内核,那么在保存现场以及传参方面究竟是怎么实现的呢?
thinking 4.1
- 内核在保存现场的时候是如何避免破坏通用寄存器的?
用
SAVE_ALL存下来了,等调用结束,再用RESTORE_SOME恢复
- 系统陷入内核调用后可以直接从当时的 \(a0-a3\) 参数寄存器中得到用户调用
msyscall留下的信息吗?当然可以,因为中间并不会被修改
- 我们是怎么做到让
sys开头的函数“认为”我们提供了和用户调用msyscall时同样 的参数的?在
user/lib/syscall_all.c中调用syscall_*时,我们将函数的前四个参数存入 \(a0 - a3\) ,剩下两个参数存进栈帧,其中第一个参数是类型编号,后五个参数是之后要用到的参数;之后调用msyscall、分发异常至handle_sys、调用do_syscall,此时我们取出 \(a1 ~ a3\) 以及栈帧中的两个数,得到的就是最开始的后五个参数。
- 内核处理系统调用的过程对
Trapframe做了哪些更改?这种修改对应的用户态的变 化是什么?修改:使
EPC加4;作用:返回后执行下一条指令
基础系统调用函数
首先要实现 kern/env.c 中的 envid2env 函数。 实现通过一个进程的 id 获取该进程控制块的功能。这个函数作为后续的基础函数会经常用到。不妨对比 envid 的产生方式:
u_int mkenvid(struct Env *e) {
static u_int i = 0;
return ((++i) << (1 + LOG2NENV)) | (e - envs);
}
#define LOG2NENV 10
#define NENV (1 << LOG2NENV)
#define ENVX(envid) ((envid) & (NENV - 1))
可知,只需要取低位即可对应出进程控制块对应的序号 e - envs。
exercise 4.3
int envid2env(u_int envid, struct Env **penv, int checkperm) {
struct Env *e = NULL;
/* Step 1: Assign value to 'e' using 'envid'. */
/* Hint:
* If envid is zero, set 'penv' to 'curenv'.
* You may want to use 'ENVX'.
*/
/* Exercise 4.3: Your code here. (1/2) */
if(envid == 0) {
*penv = curenv;
return 0;
}
e = &envs[ENVX(envid)];
if (e->env_status == ENV_FREE || e->env_id != envid) {
return -E_BAD_ENV;
}
/* Step 2: Check when 'checkperm' is non-zero. */
/* Hints:
* Check whether the calling env has sufficient permissions to manipulate the
* specified env, i.e. 'e' is either 'curenv' or its immediate child.
* If violated, return '-E_BAD_ENV'.
*/
/* Exercise 4.3: Your code here. (2/2) */
if(checkperm) {
if(e->env_id != curenv->env_id && e->env_parent_id != curenv->env_id) {
*penv = 0;
return -E_BAD_ENV;
}
}
/* Step 3: Assign 'e' to '*penv'. */
*penv = e;
return 0;
}
当然需要提醒的一点:
- 当传入参数为
0时默认处理的是当前进程curenv的信息,这一点后续经常用到。 - 还需要对是否是当前进程及其子进程进行判断,有一个权限位1
thinking 4.2
思考
envid2env函数: 为什么envid2env中需要判断e->env_id != envid的情况?如果没有这步判断会发生什么情况?这是因为
envid2env是用户调用的,无法保证合法,需要进行判断
先看向 sys_mem_alloc。这个函数的主要功能是分配内存,简单的说,用户程序可以通过这个系统调用给该程序所允许的虚拟内存空间显式地分配实际的物理内存。对于我们程 序员的视角而言,是我们编写的程序在内存中申请了一片空间;而对于操作系统内核来说,是 一个进程请求将其运行空间中的某段地址与实际物理内存进行映射,从而可以通过该虚拟页面 来对物理内存进行存取访问。内核将通过传入的进程标识符参数 (envid) 来确定发出请求的进程。
exercise 4.4
int sys_mem_alloc(u_int envid, u_int va, u_int perm) {
struct Env *env;
struct Page *pp;
/* Step 1: Check if 'va' is a legal user virtual address using 'is_illegal_va'. */
/* Exercise 4.4: Your code here. (1/3) */
if(is_illegal_va(va))
return -E_INVAL;
/* Step 2: Convert the envid to its corresponding 'struct Env *' using 'envid2env'. */
/* Hint: **Always** validate the permission in syscalls! */
/* Exercise 4.4: Your code here. (2/3) */
try(envid2env(envid, &env, 1));
/* Step 3: Allocate a physical page using 'page_alloc'. */
/* Exercise 4.4: Your code here. (3/3) */
try(page_alloc(&pp));
/* Step 4: Map the allocated page at 'va' with permission 'perm' using 'page_insert'. */
return page_insert(env->env_pgdir, env->env_asid, pp, va, perm);
}
这里的 Always 很有趣。这意味着映射也只能是跟当前进程以及子进程映射啥的,不然很容易随随便便就把别的进程清零乱搞掉。
之后是 sys_mem_map。这个函数的参数很多,但是意义很直接:将源进程地址空间中的相 应内存映射到目标进程的相应地址空间的相应虚拟内存中去。换句话说,此时两者共享一页物理内存。那么具体的逻辑为:首先找到需要操作的两个进程,其次获取源进程的虚拟页面对应 的实际物理页面,最后将该物理页面与目标进程的相应地址完成映射即可。
exercise 4.5
// src -> source dst -> destination
int sys_mem_map(u_int srcid, u_int srcva, u_int dstid, u_int dstva, u_int perm) {
struct Env *srcenv;
struct Env *dstenv;
struct Page *pp;
/* Step 1: Check if 'srcva' and 'dstva' are legal user virtual addresses using
* 'is_illegal_va'. */
/* Exercise 4.5: Your code here. (1/4) */
if(is_illegal_va(srcva) || is_illegal_va(dstva))
return -E_INVAL;
/* Step 2: Convert the 'srcid' to its corresponding 'struct Env *' using 'envid2env'. */
/* Exercise 4.5: Your code here. (2/4) */
try(envid2env(srcid, &srcenv, 1));
/* Step 3: Convert the 'dstid' to its corresponding 'struct Env *' using 'envid2env'. */
/* Exercise 4.5: Your code here. (3/4) */
try(envid2env(dstid, &dstenv, 1));
/* Step 4: Find the physical page mapped at 'srcva' in the address space of 'srcid'. */
/* Return -E_INVAL if 'srcva' is not mapped. */
/* Exercise 4.5: Your code here. (4/4) */
Pte *pte;
pp = page_lookup(srcenv->env_pgdir, srcva, &pte);
if(pp == NULL) return -E_INVAL;
/* Step 5: Map the physical page at 'dstva' in the address space of 'dstid'. */
return page_insert(dstenv->env_pgdir, dstenv->env_asid, pp, dstva, perm);
}
关于内存还有一个函数:sys_mem_unmap,这个系统调用的功能是解除某个进程地址空间虚 拟内存和物理内存之间的映射关系。
exercise 4.6
int sys_mem_unmap(u_int envid, u_int va) {
struct Env *e;
/* Step 1: Check if 'va' is a legal user virtual address using 'is_illegal_va'. */
/* Exercise 4.6: Your code here. (1/2) */
if(is_illegal_va(va))
return -E_INVAL;
/* Step 2: Convert the envid to its corresponding 'struct Env *' using 'envid2env'. */
/* Exercise 4.6: Your code here. (2/2) */
try(envid2env(envid, &e, 1));
/* Step 3: Unmap the physical page at 'va' in the address space of 'envid'. */
page_remove(e->env_pgdir, e->env_asid, va);
return 0;
}
至此常用的系统调用已全部完成。
IPC 通信机制
所谓通信,最直观的一种理解就是交换数据。假如我们能够将让一个进程有能力将数据传递给另一个进程,那么进程之间自然具有了相互通信的能力。由于这些进程的地址空间之间是相互独立的,要想传递数据,我们就需要想办法把一个地址空间中的东西传给另一个地址空间。
而 IPC 机制的实现使得我们系统中的进程之间拥有了相互传递消息的能力,为后续实现 fork、文件系统服务、管道和 shell 均有着极大的帮助。这里指导书写的非常详细,不多作赘述。
exercise 4.8
// 接受方更改信息
int sys_ipc_recv(u_int dstva) {
/* Step 1: Check if 'dstva' is either zero or a legal address. */
if (dstva != 0 && is_illegal_va(dstva)) {
return -E_INVAL;
}
/* Step 2: Set 'curenv->env_ipc_recving' to 1. */
/* Exercise 4.8: Your code here. (1/8) */
curenv->env_ipc_recving = 1; // 准备接受信息
/* Step 3: Set the value of 'curenv->env_ipc_dstva'. */
/* Exercise 4.8: Your code here. (2/8) */
curenv->env_ipc_dstva = dstva; // 设置与哪个页面完成映射
/* Step 4: Set the status of 'curenv' to 'ENV_NOT_RUNNABLE' and remove it from
* 'env_sched_list'. */
/* Exercise 4.8: Your code here. (3/8) */
// 修改进程信息
curenv->env_status = ENV_NOT_RUNNABLE;
TAILQ_REMOVE(&env_sched_list, curenv, env_sched_link);
/* Step 5: Give up the CPU and block until a message is received. */
((struct Trapframe *)KSTACKTOP - 1)->regs[2] = 0;
schedule(1);
}
// 发送方,注意是 **try_send**
int sys_ipc_try_send(u_int envid, u_int value, u_int srcva, u_int perm) {
struct Env *e;
struct Page *p;
/* Step 1: Check if 'srcva' is either zero or a legal address. */
/* Exercise 4.8: Your code here. (4/8) */
// 首先检查地址是否合法
if(srcva != 0 && is_illegal_va(srcva))
return -E_INVAL;
/* Step 2: Convert 'envid' to 'struct Env *e'. */
/* This is the only syscall where the 'envid2env' should be used with 'checkperm' UNSET,
* because the target env is not restricted to 'curenv''s children. */
/* Exercise 4.8: Your code here. (5/8) */
// 找到对应进程号
try(envid2env(envid, &e, 0));
/* Step 3: Check if the target is waiting for a message. */
/* Exercise 4.8: Your code here. (6/8) */
// 检查接收方是否做好准备
if(e->env_ipc_recving != 1)
return -E_IPC_NOT_RECV;
/* Step 4: Set the target's ipc fields. */
e->env_ipc_value = value;
e->env_ipc_from = curenv->env_id;
e->env_ipc_perm = PTE_V | perm;
e->env_ipc_recving = 0;
/* Step 5: Set the target's status to 'ENV_RUNNABLE' again and insert it to the tail of
* 'env_sched_list'. */
/* Exercise 4.8: Your code here. (7/8) */
// 传输完成后修改接受方的信息
e->env_status = ENV_RUNNABLE;
TAILQ_INSERT_TAIL(&env_sched_list, e, env_sched_link);
/* Step 6: If 'srcva' is not zero, map the page at 'srcva' in 'curenv' to 'e->env_ipc_dstva'
* in 'e'. */
/* Return -E_INVAL if 'srcva' is not zero and not mapped in 'curenv'. */
if (srcva != 0) {
/* Exercise 4.8: Your code here. (8/8) */
// 将其映射后加入即可
p = page_lookup(curenv->env_pgdir, srcva, NULL);
if(p == NULL || is_illegal_va(e->env_ipc_dstva))
return -E_INVAL;
try(page_insert(e->env_pgdir, e->env_asid, p, e->env_ipc_dstva, perm));
}
return 0;
}
值得一提的是,由于在我们的用户程序中,会大量使用 srcva 为 0 的调用来表示只传 value 值,而不需要传递物理页面,换句话说,当 srcva 不为 0 时,我们才建立两个进程的页面映射关系。因此在编写相关函数时也需要注意此种情况。
thinking 4.3
思考下面的问题,并对这个问题谈谈你的理解:请回顾
kern/env.c文件中mkenvid()函数的实现,该函数不会返回 0,请结合系统调用和IPC部分的实现与envid2env()函数的行为进行解释。
对于
envid2env:由于mkenvid不可能返回 0,所以我们可以通过判断envid是否为0来选择返回对应id的Env还是curenv对于
IPC:如果传入的mkenvid为 0,也就是不合法,我们最终达到的通信效果是自己传给自己,就不会产生其它影响
fork 进程创建
如图是使用 fork 创建进程的流程图:

fork 实现的是从父进程新生成一个子进程出来,其中子进程拥有与父进程相同的内存空间。但是如果每次都复制拷贝的话过于占空间,因此出现了一种叫做 COW(copy on write) 写时复制的算法。只有在出现修改的时候才会新建一块页表存储进去。具体函数逻辑如上图。让我们从最一开始慢慢分析其如何实现的、
fork() 函数实现
首先来看这个函数的整体逻辑:
exercise 4.15
int fork(void) {
u_int child;
u_int i;
extern volatile struct Env *env;
/* Step 1: Set our TLB Mod user exception entry to 'cow_entry' if not done yet. */
if (env->env_user_tlb_mod_entry != (u_int)cow_entry) {
try(syscall_set_tlb_mod_entry(0, cow_entry));
}
/* Step 2: Create a child env that's not ready to be scheduled. */
// Hint: 'env' should always point to the current env itself, so we should fix it to the
// correct value.
child = syscall_exofork();
//debugf("child = %d\n", child);
if (child == 0) {
env = envs + ENVX(syscall_getenvid());
return 0;
}
/* Step 3: Map all mapped pages below 'USTACKTOP' into the child's address space. */
// Hint: You should use 'duppage'.
/* Exercise 4.15: Your code here. (1/2) */
for(i = 0; i < VPN(USTACKTOP); i++)
if((vpd[i >> 10] & PTE_V) && (vpt[i] & PTE_V))
duppage(child, i);
/* Step 4: Set up the child's tlb mod handler and set child's 'env_status' to
* 'ENV_RUNNABLE'. */
/* Hint:
* You may use 'syscall_set_tlb_mod_entry' and 'syscall_set_env_status'
* Child's TLB Mod user exception entry should handle COW, so set it to 'cow_entry'
*/
/* Exercise 4.15: Your code here. (2/2) */
try(syscall_set_tlb_mod_entry(child, cow_entry));
try(syscall_set_env_status(child, ENV_RUNNABLE));
return child;
}
其思路大致分为如下几步:
- 首先通过系统调用函数
syscall_set_tlb_mod_entry对父进程的内存实现COW的标记。 - 然后调用函数
syscall_exofork实现子进程的新建,之后父子进程分别独立运行。 - 之后在父子进程之间建立映射关系,使用到的函数是
duppage。 - 最后修改子进程的状态,以及设置为
COW的标记
thinking 4.4
Thinking 4.4 关于 fork 函数的两个返回值,下面说法正确的是:
A、fork 在父进程中被调用两次,产生两个返回值
B、fork 在两个进程中分别被调用一次,产生两个不同的返回值
C、fork 只在父进程中被调用了一次,在两个进程中各产生一个返回值
D、fork 只在子进程中被调用了一次,在两个进程中各产生一个返回值
答案选: C
接下来逐一分析每一个细节的实现原理。
写时复制技术
我们要讲COW 技术给到某一个进程,因此调用如下函数实现对函数体的连接:
exercise 4.12
int sys_set_tlb_mod_entry(u_int envid, u_int func) {
struct Env *env;
/* Step 1: Convert the envid to its corresponding 'struct Env *' using 'envid2env'. */
/* Exercise 4.12: Your code here. (1/2) */
try(envid2env(envid, &env, 1));
/* Step 2: Set its 'env_user_tlb_mod_entry' to 'func'. */
/* Exercise 4.12: Your code here. (2/2) */
env->env_user_tlb_mod_entry = func;
return 0;
}
首先找到对应的进程控制块后,将其对应的函数指针给到对应位置即可。而这个的实现位于内核态。
thinking 4.8
在用户态处理页写入异常,相比于在内核态处理有什么优势?
满足“微内核”要求。即使程序出现什么问题,也不会影响到内核
那么写时复制对应的这个函数指针到底是哪里呢,位于 ./user/lib/fork.c 当中有:
exercise 4.13
static void __attribute__((noreturn)) cow_entry(struct Trapframe *tf) {
u_int va = tf->cp0_badvaddr;
u_int perm;
/* Step 1: Find the 'perm' in which the faulting address 'va' is mapped. */
/* Hint: Use 'vpt' and 'VPN' to find the page table entry. If the 'perm' doesn't have
* 'PTE_COW', launch a 'user_panic'. */
/* Exercise 4.13: Your code here. (1/6) */
perm = ((Pte*)(vpt))[VPN(va)] & 0xfff;
if(!(perm & PTE_COW))
user_panic("cow_entry: page not shared");
/* Step 2: Remove 'PTE_COW' from the 'perm', and add 'PTE_D' to it. */
/* Exercise 4.13: Your code here. (2/6) */
perm = (perm & ~(PTE_COW)) | PTE_D;
/* Step 3: Allocate a new page at 'UCOW'. */
/* Exercise 4.13: Your code here. (3/6) */
try(syscall_mem_alloc(0, (void*)UCOW, perm));
/* Step 4: Copy the content of the faulting page at 'va' to 'UCOW'. */
/* Hint: 'va' may not be aligned to a page! */
/* Exercise 4.13: Your code here. (4/6) */
memcpy((void*)UCOW, (void*)ROUNDDOWN(va, BY2PG), BY2PG);
// Step 5: Map the page at 'UCOW' to 'va' with the new 'perm'.
/* Exercise 4.13: Your code here. (5/6) */
try(syscall_mem_map(0,(void *)UCOW, 0, va, perm));
// Step 6: Unmap the page at 'UCOW'.
/* Exercise 4.13: Your code here. (6/6) */
try(syscall_mem_unmap(0, UCOW));
// Step 7: Return to the faulting routine.
int r = syscall_set_trapframe(0, tf);
user_panic("syscall_set_trapframe returned %d", r);
}
这个函数实现了写时复制的具体功能。逻辑为:
- 根据
vpt中va所在页的页表项,判断其标志位是否包含PTE_COW,是则进行下一步,否 则调用user_panic()报错。 - 分配一个新的临时物理页到临时地址
UCOW,使用memcpy将va页的数据拷贝到刚刚分配 的页中。 - 将发生页写入异常的地址
va映射到临时页面上,注意设定好对应的页面标志位(即去除PTE_COW并恢复PTE_D),然后解除临时地址UCOW的内存映射。
当然只有真正发生了 COW 的时候才会执行这段逻辑,在 fork 中相当于只是指明了这段函数的运行入口。
thinking 4.9
Thinking 4.9 请思考并回答以下几个问题:
- 为什么需要将 syscall_set_tlb_mod_entry 的调用放置在 syscall_exofork 之前?
因为
syscall_exofork可能也需要处理这个异常
- 如果放置在写时复制保护机制完成之后会有怎样的效果?
给
env_user_tlb_mod赋值时发生缺页中断,但是中断处理还未设置
那么接来下是对子进程的创建函数:
exercise 4.9
int sys_exofork(void) {
struct Env *e;
/* Step 1: Allocate a new env using 'env_alloc'. */
/* Exercise 4.9: Your code here. (1/4) */
try(env_alloc(&e, curenv->env_id));
/* Step 2: Copy the current Trapframe below 'KSTACKTOP' to the new env's 'env_tf'. */
/* Exercise 4.9: Your code here. (2/4) */
e->env_tf = *((struct Trapframe*)KSTACKTOP - 1);
/* Step 3: Set the new env's 'env_tf.regs[2]' to 0 to indicate the return value in child. */
/* Exercise 4.9: Your code here. (3/4) */
e->env_tf.regs[2] = 0;
/* Step 4: Set up the new env's 'env_status' and 'env_pri'. */
/* Exercise 4.9: Your code here. (4/4) */
e->env_status = ENV_NOT_RUNNABLE;
e->env_pri = curenv->env_pri;
return e->env_id;
}
新申请一块进程控制块之后,将其对应的控制结构 Trapframe 信息拷贝进去,然后设置返回值为0,更新各个状态即可。这里因为还没有创建完成因此状态先设置成 ENV_NOT_RUNNABLE,在 fork() 的最后才会修改。
创建好之后下一步的 fork() 将实现不同的路径。对于子进程而言,走这一步后退出:
if (child == 0) {
env = envs + ENVX(syscall_getenvid());
return 0;
}
// ./user/lib/libos.c
volatile struct Env *env;
extern int main(int, char **);
void libmain(int argc, char **argv) {
// set env to point at our env structure in envs[].
env = &envs[ENVX(syscall_getenvid())];
// call user main routine
main(argc, argv);
// exit gracefully
exit();
}
这一段与 ./user/lib/libos.c 相连,相当于实现了对进程控制块的控制,其作用类似于内核态的 curenv ,但是目前没啥用QAQ
对于父进程,此时需要拷贝链接对应内存咯,也即对应这一段代码:
/* Step 3: Map all mapped pages below 'USTACKTOP' into the child's address space. */
// Hint: You should use 'duppage'.
/* Exercise 4.15: Your code here. (1/2) */
for(i = 0; i < VPN(USTACKTOP); i++)
if((vpd[i >> 10] & PTE_V) && (vpt[i] & PTE_V))
duppage(child, i);
对于用户空间而言,其可用的用户虚拟地址为 0 - USTACKTOP 这么大,回顾之前虚拟地址的结构,其是由 10 + 10 + 12 的结构组成。使用 VPN 之后相当于右移 12 位,得知一共有多少页,然后分别对页目录项和页表项的有效性进行检查,如果通过了则进行映射。
thinking 4.6
Thinking 4.6 在遍历地址空间存取页表项时你需要使用到 vpd 和 vpt 这两个指针,请参 考 user/include/lib.h 中的相关定义,思考并回答这几个问题:
- vpt 和 vpd 的作用是什么?怎样使用它们?
vpt :页表起始地址 vpd :页目录起始地址
可以直接用数组的形式使用,如 (vpd)[va>>22] , (vpt)[va>>12]
- 从实现的角度谈一下为什么进程能够通过这种方式来存取自身的页表?
用宏定义映射了用户态页目录、页表的基地址,再通过偏移读取自身页表
- 它们是如何体现自映射设计的?
vpd的地址在UVPT和UVPT + PDMAP之间,说明将页目录映射到了某一页表位置,也就是自映射
- 进程能够通过这种方式来修改自己的页表项吗?
不能,这部分对用户态来说是只读的
然后是每一页具体的映射函数 duppage 的实现:
exercise 4.10
static void duppage(u_int envid, u_int vpn) {
int r;
u_int addr;
u_int perm;
/* Step 1: Get the permission of the page. */
/* Hint: Use 'vpt' to find the page table entry. */
/* Exercise 4.10: Your code here. (1/2) */
perm = ((Pte*)(vpt))[vpn] & 0xfff;
addr = vpn * BY2PG;
/* Step 2: If the page is writable, and not shared with children, and not marked as COW yet,
* then map it as copy-on-write, both in the parent (0) and the child (envid). */
/* Hint: The page should be first mapped to the child before remapped in the parent. (Why?)
*/
/* Exercise 4.10: Your code here. (2/2) */
r = 0;
if((perm & PTE_D) && !(perm & PTE_LIBRARY)) {
perm = (perm & (~PTE_D)) | PTE_COW; // 修改权限位
r = 1;
}
syscall_mem_map(0, addr, envid, addr, perm); // 建立映射
if(r)
syscall_mem_map(0, addr, 0, addr, perm); // 改 perm
}
不难发现,这里的 envid 传入了 0 还是蛮常见的,与前文成功呼应。
thinking 4.5
我们并不应该对所有的用户空间页都使用
duppage进行映射。那么究竟哪 些用户空间页应该映射,哪些不应该呢?请结合kern/env.c中env_init函数进行的页面映射、include/mmu.h里的内存布局图以及本章的后续描述进行思考。
UXSTACKTOP以上的部分是系统空间,不映射
UXSTACKTOP-BY2PG到UXSTACKTOP之间是异常栈,写时复制的时候还得用,所以不映射
USTACKTOP到UXSTACKTOP-BY2PG中间没有内容,不映射
UTEXT到UXSTACKTOP-BY2PG中,可写且不共享的内容都需要被映射
在程序运行过程中,如果发现需要实行复制 ,此时会触发一个异常称之为页写入异常,也即位于 kern/tlbex.c 中:
exercise 4.11
void do_tlb_mod(struct Trapframe *tf) {
struct Trapframe tmp_tf = *tf;
if (tf->regs[29] < USTACKTOP || tf->regs[29] >= UXSTACKTOP) {
tf->regs[29] = UXSTACKTOP;
}
tf->regs[29] -= sizeof(struct Trapframe);
*(struct Trapframe *)tf->regs[29] = tmp_tf;
if (curenv->env_user_tlb_mod_entry) {
tf->regs[4] = tf->regs[29];
tf->regs[29] -= sizeof(tf->regs[4]);
// Hint: Set 'cp0_epc' in the context 'tf' to 'curenv->env_user_tlb_mod_entry'.
/* Exercise 4.11: Your code here. */
tf->cp0_epc = curenv->env_user_tlb_mod_entry;
} else {
panic("TLB Mod but no user handler registered");
}
}
负责将当前现场保存在异常处理栈中,并设置 a0 和 EPC 寄存器的值,使得从异常恢复后能够以异常处理栈中保存的现场(Trapframe)为参数,跳转到 env_user_tlb_mod_entry 域存储的用户异常处理函数的地址。
thinking 4.7
在
do_tlb_mod函数中,你可能注意到了一个向异常处理栈复制Trapframe运行现场的过程,请思考并回答这几个问题:
- 这里实现了一个支持类似于“异常重入”的机制,而在什么时候会出现这种“异常重入” ?
在写时复制、发生缺页异常时,可能再次发生缺页异常,从而“异常重入
- 内核为什么需要将异常的现场
Trapframe复制到用户空间?首先,我们的
MOS操作系统中,对缺页错误的处理是由用户进程完成的,用户进程在处理过程中需要读取Trapframe的内容;另外,中断结束后恢复现场也要用到
Trapframe
最后就是修改子进程的状态信息,需要用到函数为:
exercise 4.14
int sys_set_env_status(u_int envid, u_int status) {
struct Env *env;
/* Step 1: Check if 'status' is valid. */
/* Exercise 4.14: Your code here. (1/3) */
if(status != ENV_RUNNABLE && status != ENV_NOT_RUNNABLE)
return -E_INVAL;
/* Step 2: Convert the envid to its corresponding 'struct Env *' using 'envid2env'. */
/* Exercise 4.14: Your code here. (2/3) */
try(envid2env(envid, &env, 1));
/* Step 3: Update 'env_sched_list' if the 'env_status' of 'env' is being changed. */
/* Exercise 4.14: Your code here. (3/3) */
if(status == ENV_NOT_RUNNABLE && env->env_status != ENV_NOT_RUNNABLE)
TAILQ_REMOVE(&env_sched_list, env, env_sched_link);
if(status == ENV_RUNNABLE && env->env_status != ENV_RUNNABLE)
TAILQ_INSERT_TAIL(&env_sched_list, env, env_sched_link);
/* Step 4: Set the 'env_status' of 'env'. */
env->env_status = status;
return 0;
}
进行修改,并判断其如果不一致则加入或删除调度队列。
之后就维护好了一个子进程,可以独立的运行啦,而我们本次的实验到这里也正式结束。


浙公网安备 33010602011771号