memcached源码分析-----slab automove和slab rebalance

 转载请注明出处:http://blog.csdn.net/luotuo44/article/details/43015129

 

 

需求:

 

 

        考虑这样的一个情景:在一开始,由于业务原因向memcached存储大量长度为1KB的数据,也就是说memcached服务器进程里面有很多大小为1KB的item。现在由于业务调整需要存储大量10KB的数据,并且很少使用1KB的那些数据了。由于数据越来越多,内存开始吃紧。大小为10KB的那些item频繁访问,并且由于内存不够需要使用LRU淘汰一些10KB的item。

        对于上面的情景,会不会觉得大量1KB的item实在太浪费了。由于很少访问这些item,所以即使它们超时过期了,还是会占据着哈希表和LRU队列。LRU队列还好,不同大小的item使用不同的LRU队列。但对于哈希表来说大量的僵尸item会增加哈希冲突的可能性,并且在迁移哈希表的时候也浪费时间。有没有办法干掉这些item?使用LRU爬虫+lru_crawler命令是可以强制干掉这些僵尸item。但干掉这些僵尸item后,它们占据的内存是归还到1KB的那些slab分配器中。1KB的slab分配器不会为10KB的item分配内存。所以还是功亏一篑。

 

        那有没有别的办法呢?是有的。memcached提供的slab automove 和 rebalance两个东西就是完成这个功能的。在默认情况下,memcached不启动这个功能,所以要想使用这个功能必须在启动memcached的时候加上参数-o slab_reassign。之后就可以在客户端发送命令slabsreassign <source class> <dest class>,手动将source class的内存页分给dest class。后文会把这个工作称为内存页重分配。而命令slabs automove则是让memcached自动检测是否需要进行内存页重分配,如果需要的话就自动去操作,这样一切都不需要人工的干预。

        如果在启动memcached的时候使用了参数-o slab_reassign,那么就会把settings.slab_reassign赋值为true(该变量的默认值为false)。还记得《slab内存分配器》说到的每一个内存页的大小吗?在do_slabs_newslab函数中,一个内存页的大小会根据settings.slab_reassign是否为true而不同。

[cpp] view plain copy
 
 在CODE上查看代码片派生到我的代码片
  1. static int do_slabs_newslab(const unsigned int id) {  
  2.     slabclass_t *p = &slabclass[id];  
  3.     //settings.slab_reassign的默认值为false  
  4.     int len = settings.slab_reassign ? settings.item_size_max  
  5.         : p->size * p->perslab;  
  6.   
  7.     //len就是一个内存页的大小  
  8.     ...  
  9. }  

 

        当settings.slab_reassign为true,也就是启动rebalance功能的时候,slabclass数组中所有slabclass_t的内存页都是一样大的,等于settings.item_size_max(默认为1MB)。这样做的好处就是在需要将一个内存页从某一个slabclass_t强抢给另外一个slabclass_t时,比较好处理。不然的话,slabclass[i]从slabclass[j] 抢到的一个内存页可以切分为n个item,而从slabclass[k]抢到的一个内存页却切分为m个item,而本身的一个内存页有s个item。这样的话是相当混乱的。假如毕竟统一了内存页大小,那么无论从哪里抢到的内存页都是切分成一样多的item个数。

 

 

启动和终止rebalance:

 

        main函数会调用start_slab_maintenance_thread函数启动rebalance线程和automove线程。main函数是在settings.slab_reassign为true时才会调用的。

[cpp] view plain copy
 
 在CODE上查看代码片派生到我的代码片
  1. //slabs.c文件  
  2. static pthread_cond_t maintenance_cond = PTHREAD_COND_INITIALIZER;  
  3. static pthread_cond_t slab_rebalance_cond = PTHREAD_COND_INITIALIZER;  
  4. static volatile int do_run_slab_thread = 1;  
  5. static volatile int do_run_slab_rebalance_thread = 1;  
  6.   
  7. #define DEFAULT_SLAB_BULK_CHECK 1  
  8. int slab_bulk_check = DEFAULT_SLAB_BULK_CHECK;  
  9.   
  10. static pthread_mutex_t slabs_lock = PTHREAD_MUTEX_INITIALIZER;  
  11. static pthread_mutex_t slabs_rebalance_lock = PTHREAD_MUTEX_INITIALIZER;  
  12.   
  13. static pthread_t maintenance_tid;  
  14. static pthread_t rebalance_tid;  
  15.   
  16.   
  17.   
  18. //由main函数调用,如果settings.slab_reassign为false将不会调用本函数(默认是false)  
  19. int start_slab_maintenance_thread(void) {  
  20.     int ret;  
  21.     slab_rebalance_signal = 0;  
  22.     slab_rebal.slab_start = NULL;  
  23.     char *env = getenv("MEMCACHED_SLAB_BULK_CHECK");  
  24.     if (env != NULL) {  
  25.         slab_bulk_check = atoi(env);  
  26.         if (slab_bulk_check == 0) {  
  27.             slab_bulk_check = DEFAULT_SLAB_BULK_CHECK;  
  28.         }  
  29.     }  
  30.   
  31.     if (pthread_cond_init(&slab_rebalance_cond, NULL) != 0) {  
  32.         fprintf(stderr, "Can't intiialize rebalance condition\n");  
  33.         return -1;  
  34.     }  
  35.     pthread_mutex_init(&slabs_rebalance_lock, NULL);  
  36.   
  37.     if ((ret = pthread_create(&maintenance_tid, NULL,  
  38.                               slab_maintenance_thread, NULL)) != 0) {  
  39.         fprintf(stderr, "Can't create slab maint thread: %s\n", strerror(ret));  
  40.         return -1;  
  41.     }  
  42.     if ((ret = pthread_create(&rebalance_tid, NULL,  
  43.                               slab_rebalance_thread, NULL)) != 0) {  
  44.         fprintf(stderr, "Can't create rebal thread: %s\n", strerror(ret));  
  45.         return -1;  
  46.     }  
  47.     return 0;  
  48. }  
  49.   
  50. void stop_slab_maintenance_thread(void) {  
  51.     mutex_lock(&cache_lock);  
  52.     do_run_slab_thread = 0;  
  53.     do_run_slab_rebalance_thread = 0;  
  54.     pthread_cond_signal(&maintenance_cond);  
  55.     pthread_mutex_unlock(&cache_lock);  
  56.   
  57.     /* Wait for the maintenance thread to stop */  
  58.     pthread_join(maintenance_tid, NULL);  
  59.     pthread_join(rebalance_tid, NULL);  
  60. }  

 

        要注意的是,start_slab_maintenance_thread函数启动了两个线程:rebalance线程和automove线程。automove线程会自动检测是否需要进行内存页重分配。如果检测到需要重分配,那么就会叫rebalance线程执行这个内存页重分配工作。

        默认情况下是不开启自动检测功能的,即使在启动memcached的时候加入了-o slab_reassign参数。自动检测功能由全局变量settings.slab_automove控制(默认值为0,0就是不开启)。如果要开启可以在启动memcached的时候加入slab_automove选项,并将其参数数设置为1。比如命令$memcached -o slab_reassign,slab_automove=1就开启了自动检测功能。当然也是可以在启动memcached后通过客户端命令启动automove功能,使用命令slabsautomove <0|1>。其中0表示关闭automove,1表示开启automove。客户端的这个命令只是简单地设置settings.slab_automove的值,不做其他任何工作。

 

 

automove线程:

 

 

item状态记录仪:

        由于rebalance线程启动后就会由于等待条件变量而进入休眠状态,等待别人给它内存页重分配任务。所以我们先来看一下automove线程。

        automove线程要进行自动检测,检测就需要一些实时数据进行分析。然后得出结论:哪个slabclass_t需要更多的内存,哪个又不需要。automove线程通过全局变量itemstats收集item的各种数据。下面看一下itemstats变量以及它的类型定义。

[cpp] view plain copy
 
 在CODE上查看代码片派生到我的代码片
  1. //items.c文件  
  2. typedef struct {  
  3.     uint64_t evicted;//因为LRU踢了多少个item  
  4.     //即使一个item的exptime设置为0,也是会被踢的  
  5.     uint64_t evicted_nonzero;//被踢的item中,超时时间(exptime)不为0的item数  
  6.   
  7.     //最后一次踢item时,被踢的item已经过期多久了  
  8.     //itemstats[id].evicted_time = current_time - search->time;  
  9.     rel_time_t evicted_time;  
  10.   
  11.       
  12.     uint64_t reclaimed;//在申请item时,发现过期并回收的item数量  
  13.     uint64_t outofmemory;//为item申请内存,失败的次数  
  14.     uint64_t tailrepairs;//需要修复的item数量(除非worker线程有问题否则一般为0)  
  15.       
  16.     //直到被超时删除时都还没被访问过的item数量  
  17.     uint64_t expired_unfetched;  
  18.     //直到被LRU踢出时都还没有被访问过的item数量  
  19.     uint64_t evicted_unfetched;  
  20.       
  21.     uint64_t crawler_reclaimed;//被LRU爬虫发现的过期item数量  
  22.   
  23.     //申请item而搜索LRU队列时,被其他worker线程引用的item数量  
  24.     uint64_t lrutail_reflocked;  
  25. } itemstats_t;  
  26.   
  27. #define POWER_LARGEST  200  
  28. #define LARGEST_ID POWER_LARGEST  
  29. static itemstats_t itemstats[LARGEST_ID];  

 

        注意上面代码是在items.c文件的,并且全局变量itemstats是static类型。itemstats变量是一个数组,它是和slabclass数组一一对应的。itemstats数组的元素负责收集slabclass数组中对应元素的信息。itemstats_t结构体虽然提供了很多成员,可以收集很多信息,但automove线程只用到第一个成员evicted。automove线程需要知道每一个尺寸的item的被踢情况,然后判断哪一类item资源紧缺,哪一类item资源又过剩。

        itemstats广泛分布在items.c文件的多个函数中(主要是为了能收集各种数据),所以这里就不给出itemstats的具体收集实现了。当然由于evicted是重要的而且只在一个函数出现,就贴出evicted的收集代码吧。

[cpp] view plain copy
 
 在CODE上查看代码片派生到我的代码片
  1. item *do_item_alloc(char *key, const size_t nkey, const int flags,  
  2.                     const rel_time_t exptime, const int nbytes,  
  3.                     const uint32_t cur_hv) {  
  4.     item *it = NULL;  
  5.   
  6.     int tries = 5;  
  7.     item *search;  
  8.     item *next_it;  
  9.     rel_time_t oldest_live = settings.oldest_live;  
  10.   
  11.     search = tails[id];  
  12.     for (; tries > 0 && search != NULL; tries--, search=next_it) {  
  13.         /* we might relink search mid-loop, so search->prev isn't reliable */  
  14.         next_it = search->prev;  
  15.   
  16.         ...  
  17.           
  18.         if ((search->exptime != 0 && search->exptime < current_time)  
  19.             || (search->time <= oldest_live && oldest_live <= current_time)) {  
  20.             ...   
  21.         } else if ((it = slabs_alloc(ntotal, id)) == NULL) {//申请内存失败  
  22.             //此刻,过期失效的item没有找到,申请内存又失败了。看来只能使用  
  23.             //LRU淘汰一个item(即使这个item并没有过期失效)  
  24.               
  25.             if (settings.evict_to_free == 0) {//设置了不进行LRU淘汰item  
  26.                 //此时只能向客户端回复错误了  
  27.                 itemstats[id].outofmemory++;  
  28.             } else {  
  29.                 itemstats[id].evicted++;//增加被踢的item数  
  30.                 itemstats[id].evicted_time = current_time - search->time;  
  31.                 //即使一个item的exptime成员设置为永不超时(0),还是会被踢的  
  32.                 if (search->exptime != 0)  
  33.                     itemstats[id].evicted_nonzero++;  
  34.                 if ((search->it_flags & ITEM_FETCHED) == 0) {  
  35.                     itemstats[id].evicted_unfetched++;  
  36.                 }  
  37.                 it = search;  
  38.   
  39.                 //一旦发现有item被踢,那么就启动内存页重分配操作  
  40.                 //这个太频繁了,不推荐                  
  41.                 if (settings.slab_automove == 2)  
  42.                     slabs_reassign(-1, id);  
  43.             }  
  44.         }  
  45.   
  46.         break;  
  47.     }  
  48.   
  49.     ...  
  50.     return it;  
  51. }  

 

        从上面的代码可以看到,如果某个item因为LRU被踢了,那么就会被记录起来。在最后还可以看到如果settings.slab_automove 等于2,那么一旦有item被踢了就调用slabs_reassign函数。slabs_reassign函数就是内存页重分配处理函数。明显一有item被踢就重分配太频繁了,所以这是不推荐的。

 

 

确定贫穷和富有item:

        现在回过来看一下automove线程的线程函数slab_maintenance_thread。

[cpp] view plain copy
 
 在CODE上查看代码片派生到我的代码片
  1. static void *slab_maintenance_thread(void *arg) {  
  2.     int src, dest;  
  3.   
  4.     while (do_run_slab_thread) {  
  5.         if (settings.slab_automove == 1) {//启动了automove功能  
  6.             if (slab_automove_decision(&src, &dest) == 1) {  
  7.                 /* Blind to the return codes. It will retry on its own */  
  8.                 slabs_reassign(src, dest);  
  9.             }  
  10.             sleep(1);  
  11.         } else {//等待用户启动automove  
  12.             /* Don't wake as often if we're not enabled. 
  13.              * This is lazier than setting up a condition right now. */  
  14.             sleep(5);  
  15.         }  
  16.     }  
  17.     return NULL;  
  18. }  

 

        可以看到如果settings.slab_automove就调用slab_automove_decision判断是否应该进行内存页重分配。返回1就说明需要重分配内存页,此时调用slabs_reassign进行处理。现在来看一下automove线程是怎么判断要不要进行内存页重分配的。

[cpp] view plain copy
 
 在CODE上查看代码片派生到我的代码片
  1. //items.c文件  
  2. void item_stats_evictions(uint64_t *evicted) {  
  3.     int i;  
  4.     mutex_lock(&cache_lock);  
  5.     for (i = 0; i < LARGEST_ID; i++) {  
  6.         evicted[i] = itemstats[i].evicted;  
  7.     }  
  8.     mutex_unlock(&cache_lock);  
  9. }  
  10.   
  11.   
  12. //slabs.c文件  
  13. //本函数选出最佳被踢选手,和最佳不被踢选手。返回1表示成功选手两位选手  
  14. //返回0表示没有选出。要同时选出两个选手才返回1。并用src参数记录最佳不  
  15. //不踢选手的id,dst记录最佳被踢选手的id  
  16. static int slab_automove_decision(int *src, int *dst) {  
  17.     static uint64_t evicted_old[POWER_LARGEST];  
  18.     static unsigned int slab_zeroes[POWER_LARGEST];  
  19.     static unsigned int slab_winner = 0;  
  20.     static unsigned int slab_wins   = 0;  
  21.     uint64_t evicted_new[POWER_LARGEST];  
  22.     uint64_t evicted_diff = 0;  
  23.     uint64_t evicted_max  = 0;  
  24.     unsigned int highest_slab = 0;  
  25.     unsigned int total_pages[POWER_LARGEST];  
  26.     int i;  
  27.     int source = 0;  
  28.     int dest = 0;  
  29.     static rel_time_t next_run;  
  30.   
  31.     /* Run less frequently than the slabmove tester. */  
  32.     //本函数的调用不能过于频繁,至少10秒调用一次  
  33.     if (current_time >= next_run) {  
  34.         next_run = current_time + 10;  
  35.     } else {  
  36.         return 0;  
  37.     }  
  38.   
  39.     //获取每一个slabclass的被踢item数  
  40.     item_stats_evictions(evicted_new);  
  41.     pthread_mutex_lock(&cache_lock);  
  42.     for (i = POWER_SMALLEST; i < power_largest; i++) {  
  43.         total_pages[i] = slabclass[i].slabs;  
  44.     }  
  45.     pthread_mutex_unlock(&cache_lock);  
  46.   
  47.     //本函数会频繁被调用,所以有次数可说。  
  48.       
  49.     /* Find a candidate source; something with zero evicts 3+ times */  
  50.     //evicted_old记录上一个时刻每一个slabclass的被踢item数  
  51.     //evicted_new则记录了现在每一个slabclass的被踢item数  
  52.     //evicted_diff则能表现某一个LRU队列被踢的频繁程度  
  53.     for (i = POWER_SMALLEST; i < power_largest; i++) {  
  54.         evicted_diff = evicted_new[i] - evicted_old[i];  
  55.         if (evicted_diff == 0 && total_pages[i] > 2) {  
  56.             //evicted_diff等于0说明这个slabclass没有item被踢,而且  
  57.             //它又占有至少两个slab。           
  58.             slab_zeroes[i]++;//增加计数  
  59.             //这个slabclass已经历经三次都没有被踢记录,说明空间多得很  
  60.             //就选你了,最佳不被踢选手  
  61.             if (source == 0 && slab_zeroes[i] >= 3)  
  62.                 source = i;  
  63.         } else {  
  64.             slab_zeroes[i] = 0;//计数清零  
  65.             if (evicted_diff > evicted_max) {  
  66.                 evicted_max = evicted_diff;  
  67.                 highest_slab = i;  
  68.             }  
  69.         }  
  70.         evicted_old[i] = evicted_new[i];  
  71.     }  
  72.   
  73.     /* Pick a valid destination */  
  74.     //选出一个slabclass,这个slabclass要连续3次都是被踢最多item的那个slabclass  
  75.     if (slab_winner != 0 && slab_winner == highest_slab) {  
  76.         slab_wins++;  
  77.         if (slab_wins >= 3)//这个slabclass已经连续三次成为最佳被踢选手了  
  78.             dest = slab_winner;  
  79.     } else {  
  80.         slab_wins = 1;//计数清零(当然这里是1)  
  81.         slab_winner = highest_slab;//本次的最佳被踢选手  
  82.     }  
  83.   
  84.     if (source && dest) {  
  85.         *src = source;  
  86.         *dst = dest;  
  87.         return 1;  
  88.     }  
  89.     return 0;  
  90. }  

 

        从上面的代码也可以看到,其实判断的方法也比较简单。从slabclass数组中选出两个选手:一个是连续三次没有被踢item了,另外一个则是连续三次都成为最佳被踢手。如果找到了满足条件的两个选手,那么返回1。此时automove线程就会调用slabs_reassign函数。

 

下达 rebalance任务:

        在贴出slabs_reassign函数前,回想一下slabs reassign命令。前面讲的都是自动检测要不要进行内存页重分配,都快要忘了还有一个手动要求内存页重分配的命令。如果客户端使用了slabs reassign命令,那么worker线程在接收到这个命令后,就会调用slabs_reassign函数,函数参数是slabs reassign命令的参数。现在自动检测和手动设置大一统了。

[cpp] view plain copy
 
 在CODE上查看代码片派生到我的代码片
  1. enum reassign_result_type {  
  2.     REASSIGN_OK=0, REASSIGN_RUNNING, REASSIGN_BADCLASS, REASSIGN_NOSPARE,  
  3.     REASSIGN_SRC_DST_SAME  
  4. };  
  5.   
  6.   
  7. enum reassign_result_type slabs_reassign(int src, int dst) {  
  8.     enum reassign_result_type ret;  
  9.     if (pthread_mutex_trylock(&slabs_rebalance_lock) != 0) {  
  10.         return REASSIGN_RUNNING;  
  11.     }  
  12.     ret = do_slabs_reassign(src, dst);  
  13.     pthread_mutex_unlock(&slabs_rebalance_lock);  
  14.     return ret;  
  15. }  
  16.   
  17.   
  18. static enum reassign_result_type do_slabs_reassign(int src, int dst) {  
  19.     if (slab_rebalance_signal != 0)  
  20.         return REASSIGN_RUNNING;  
  21.   
  22.     if (src == dst)//不能相同  
  23.         return REASSIGN_SRC_DST_SAME;  
  24.   
  25.     /* Special indicator to choose ourselves. */  
  26.     if (src == -1) {//客户端命令要求随机选出一个源slab class  
  27.         //选出一个页数大于1的slab class,并且该slab class不能是dst  
  28.         //指定的那个。如果不存在这样的slab class,那么返回-1  
  29.         src = slabs_reassign_pick_any(dst);  
  30.         /* TODO: If we end up back at -1, return a new error type */  
  31.     }  
  32.   
  33.     if (src < POWER_SMALLEST || src > power_largest ||  
  34.         dst < POWER_SMALLEST || dst > power_largest)  
  35.         return REASSIGN_BADCLASS;  
  36.   
  37.     //源slab class没有或者只有一个内存页,那么就不能分给别的slab class  
  38.     if (slabclass[src].slabs < 2)  
  39.         return REASSIGN_NOSPARE;  
  40.   
  41.     //全局变量slab_rebal  
  42.     slab_rebal.s_clsid = src;//保存源slab class  
  43.     slab_rebal.d_clsid = dst;//保存目标slab class  
  44.   
  45.     slab_rebalance_signal = 1;  
  46.     //唤醒slab_rebalance_thread函数的线程.  
  47.     //在slabs_reassign函数中已经锁上了slabs_rebalance_lock  
  48.     pthread_cond_signal(&slab_rebalance_cond);  
  49.   
  50.     return REASSIGN_OK;  
  51. }  
  52.   
  53.   
  54. //选出一个内存页数大于1的slab class,并且该slab class不能是dst  
  55. //指定的那个。如果不存在这样的slab class,那么返回-1  
  56. static int slabs_reassign_pick_any(int dst) {  
  57.     static int cur = POWER_SMALLEST - 1;  
  58.     int tries = power_largest - POWER_SMALLEST + 1;  
  59.     for (; tries > 0; tries--) {  
  60.         cur++;  
  61.         if (cur > power_largest)  
  62.             cur = POWER_SMALLEST;  
  63.         if (cur == dst)  
  64.             continue;  
  65.         if (slabclass[cur].slabs > 1) {  
  66.             return cur;  
  67.         }  
  68.     }  
  69.     return -1;  
  70. }  

 

        do_slabs_reassign会把源slab class 和目标slab class保存在全局变量slab_rebal,并且在最后会调用pthread_cond_signal唤醒rebalance线程。

 

 

rebalance线程:

 

        现在automove线程已经退出历史舞台了,rebalance线程也从沉睡中苏醒过来并登上舞台。现在来看一下rebalance线程的线程函数slab_rebalance_thread。注意:在一开始slab_rebalance_signal是等于0的,当需要进行内存页重分配就会把slab_rebalance_signal变量赋值为1。

[cpp] view plain copy
 
 在CODE上查看代码片派生到我的代码片
  1. static void *slab_rebalance_thread(void *arg) {  
  2.     int was_busy = 0;  
  3.     /* So we first pass into cond_wait with the mutex held */  
  4.     mutex_lock(&slabs_rebalance_lock);  
  5.   
  6.     while (do_run_slab_rebalance_thread) {  
  7.         if (slab_rebalance_signal == 1) {  
  8.             //标志要移动的内存页的信息,并将slab_rebalance_signal赋值为2  
  9.             //slab_rebal.done赋值为0,表示没有完成  
  10.             if (slab_rebalance_start() < 0) {//失败  
  11.                 /* Handle errors with more specifity as required. */  
  12.                 slab_rebalance_signal = 0;  
  13.             }  
  14.   
  15.             was_busy = 0;  
  16.         } else if (slab_rebalance_signal && slab_rebal.slab_start != NULL) {  
  17.             was_busy = slab_rebalance_move();//进行内存页迁移操作  
  18.         }  
  19.   
  20.         if (slab_rebal.done) {//完成内存页重分配操作  
  21.             slab_rebalance_finish();  
  22.         } else if (was_busy) {//有worker线程在使用内存页上的item  
  23.             /* Stuck waiting for some items to unlock, so slow down a bit 
  24.              * to give them a chance to free up */  
  25.             usleep(50);//休眠一会儿,等待worker线程放弃使用item,然后再次尝试  
  26.         }  
  27.   
  28.         if (slab_rebalance_signal == 0) {//一开始就在这里休眠  
  29.             /* always hold this lock while we're running */  
  30.             pthread_cond_wait(&slab_rebalance_cond, &slabs_rebalance_lock);  
  31.         }  
  32.     }  
  33.     return NULL;  
  34. }  



锁定内存页:

        函数slab_rebalance_start对要源slab class进行一些标注,当worker线程要访问源slab class的时候意识到正在内存页重分配。

 

[cpp] view plain copy
 
 在CODE上查看代码片派生到我的代码片
  1. //memcached.h文件  
  2. struct slab_rebalance {  
  3.     //记录要移动的页的信息。slab_start指向页的开始位置。slab_end指向页  
  4.     //的结束位置。slab_pos则记录当前处理的位置(item)  
  5.     void *slab_start;  
  6.     void *slab_end;  
  7.     void *slab_pos;  
  8.     int s_clsid; //源slab class的下标索引  
  9.     int d_clsid; //目标slab class的下标索引  
  10.     int busy_items; //是否worker线程在引用某个item  
  11.     uint8_t done;//是否完成了内存页移动  
  12. };  
  13. //memcached.c文件  
  14. struct slab_rebalance slab_rebal;  
  15.   
  16. //slabs.c文件  
  17. static int slab_rebalance_start(void) {  
  18.     slabclass_t *s_cls;  
  19.     int no_go = 0;  
  20.   
  21.     pthread_mutex_lock(&cache_lock);  
  22.     pthread_mutex_lock(&slabs_lock);  
  23.   
  24.     if (slab_rebal.s_clsid < POWER_SMALLEST ||  
  25.         slab_rebal.s_clsid > power_largest  ||  
  26.         slab_rebal.d_clsid < POWER_SMALLEST ||  
  27.         slab_rebal.d_clsid > power_largest  ||  
  28.         slab_rebal.s_clsid == slab_rebal.d_clsid)//非法下标索引  
  29.         no_go = -2;  
  30.   
  31.     s_cls = &slabclass[slab_rebal.s_clsid];  
  32.   
  33.     //为这个目标slab class增加一个页表项都失败,那么就  
  34.     //根本无法为之增加一个页了  
  35.     if (!grow_slab_list(slab_rebal.d_clsid)) {  
  36.         no_go = -1;  
  37.     }  
  38.   
  39.     if (s_cls->slabs < 2)//目标slab class页数太少了,无法分一个页给别人  
  40.         no_go = -3;  
  41.   
  42.     if (no_go != 0) {  
  43.         pthread_mutex_unlock(&slabs_lock);  
  44.         pthread_mutex_unlock(&cache_lock);  
  45.         return no_go; /* Should use a wrapper function... */  
  46.     }  
  47.   
  48.     //标志将源slab class的第几个内存页分给目标slab class  
  49.     //这里是默认是将第一个内存页分给目标slab class  
  50.     s_cls->killing = 1;  
  51.   
  52.     //记录要移动的页的信息。slab_start指向页的开始位置。slab_end指向页  
  53.     //的结束位置。slab_pos则记录当前处理的位置(item)  
  54.     slab_rebal.slab_start = s_cls->slab_list[s_cls->killing - 1];  
  55.     slab_rebal.slab_end   = (char *)slab_rebal.slab_start +  
  56.         (s_cls->size * s_cls->perslab);  
  57.     slab_rebal.slab_pos   = slab_rebal.slab_start;  
  58.     slab_rebal.done       = 0;  
  59.   
  60.     /* Also tells do_item_get to search for items in this slab */  
  61.     slab_rebalance_signal = 2;//要rebalance线程接下来进行内存页移动  
  62.     
  63.   
  64.     pthread_mutex_unlock(&slabs_lock);  
  65.     pthread_mutex_unlock(&cache_lock);  
  66.   
  67.     return 0;  
  68. }  

 

 

        slab_rebalance_start会将一个slab class的一个内存页标注为要移动的,此时就不能让worker线程访问这个内存页的item了。现在看一下假如worker线程刚好要访问这个内存页的一个item时会发生什么。

[cpp] view plain copy
 
 在CODE上查看代码片派生到我的代码片
  1. item *do_item_get(const char *key, const size_t nkey, const uint32_t hv) {  
  2.     item *it = assoc_find(key, nkey, hv);//assoc_find函数内部没有加锁  
  3.       
  4.     if (it != NULL) {//找到了,此时item的引用计数至少为1  
  5.         refcount_incr(&it->refcount);//线程安全地自增一  
  6.         /* Optimization for slab reassignment. prevents popular items from 
  7.          * jamming in busy wait. Can only do this here to satisfy lock order 
  8.          * of item_lock, cache_lock, slabs_lock. */  
  9.         if (slab_rebalance_signal &&  
  10.             ((void *)it >= slab_rebal.slab_start && (void *)it < slab_rebal.slab_end)) {  
  11.             //这个item刚好在要移动的内存页里面。此时不能返回这个item  
  12.             //worker线程要负责把这个item从哈希表和LRU队列中删除这个item,避免  
  13.             //后面有其他worker线程又访问这个不能使用的item  
  14.             do_item_unlink_nolock(it, hv);  
  15.             do_item_remove(it);  
  16.             it = NULL;  
  17.         }  
  18.     }  
  19.   
  20.     ...  
  21.     return it;  
  22. }  

 

 

移动(归还)item:

        现在回过头继续看rebalance线程。前面说到已经标注了源slab class的一个内存页。标注完rebalance线程就会调用slab_rebalance_move函数完成真正的内存页迁移操作。源slab class上的内存页是有item的,那么在迁移的时候怎么处理这些item呢?memcached的处理方式是很粗暴的:直接删除。如果这个item还有worker线程在使用,rebalance线程就等你一下。如果这个item没有worker线程在引用,那么即使这个item没有过期失效也将直接删除。

        因为一个内存页可能会有很多个item,所以memcached也采用分期处理的方法,每次只处理少量的item(默认为一个)。所以呢,slab_rebalance_move函数会在slab_rebalance_thread线程函数中多次调用,直到处理了所有的item。

[cpp] view plain copy
 
 在CODE上查看代码片派生到我的代码片
  1. /* refcount == 0 is safe since nobody can incr while cache_lock is held. 
  2.  * refcount != 0 is impossible since flags/etc can be modified in other 
  3.  * threads. instead, note we found a busy one and bail. logic in do_item_get 
  4.  * will prevent busy items from continuing to be busy 
  5.  */  
  6. static int slab_rebalance_move(void) {  
  7.     slabclass_t *s_cls;  
  8.     int x;  
  9.     int was_busy = 0;  
  10.     int refcount = 0;  
  11.     enum move_status status = MOVE_PASS;  
  12.   
  13.     pthread_mutex_lock(&cache_lock);  
  14.     pthread_mutex_lock(&slabs_lock);  
  15.   
  16.     s_cls = &slabclass[slab_rebal.s_clsid];  
  17.   
  18.     //会在start_slab_maintenance_thread函数中读取环境变量设置slab_bulk_check  
  19.     //默认值为1.同样这里也是采用分期处理的方案处理一个页上的多个item  
  20.     for (x = 0; x < slab_bulk_check; x++) {  
  21.         item *it = slab_rebal.slab_pos;  
  22.         status = MOVE_PASS;  
  23.         if (it->slabs_clsid != 255) {  
  24.             void *hold_lock = NULL;  
  25.             uint32_t hv = hash(ITEM_key(it), it->nkey);  
  26.             if ((hold_lock = item_trylock(hv)) == NULL) {  
  27.                 status = MOVE_LOCKED;  
  28.             } else {  
  29.                 refcount = refcount_incr(&it->refcount);  
  30.                 if (refcount == 1) { /* item is unlinked, unused */  
  31.                     //如果it_flags&ITEM_SLABBED为真,那么就说明这个item  
  32.                     //根本就没有分配出去。如果为假,那么说明这个item被分配  
  33.                     //出去了,但处于归还途中。参考do_item_get函数里面的  
  34.                     //判断语句,有slab_rebalance_signal作为判断条件的那个。  
  35.                     if (it->it_flags & ITEM_SLABBED) {//没有分配出去  
  36.                         /* remove from slab freelist */  
  37.                         if (s_cls->slots == it) {  
  38.                             s_cls->slots = it->next;  
  39.                         }  
  40.                         if (it->next) it->next->prev = it->prev;  
  41.                         if (it->prev) it->prev->next = it->next;  
  42.                         s_cls->sl_curr--;  
  43.                         status = MOVE_DONE;//这个item处理成功  
  44.                     } else {//此时还有另外一个worker线程在归还这个item  
  45.                         status = MOVE_BUSY;  
  46.                     }  
  47.                 } else if (refcount == 2) { /* item is linked but not busy */  
  48.                     //没有worker线程引用这个item  
  49.                     if ((it->it_flags & ITEM_LINKED) != 0) {  
  50.                         //直接把这个item从哈希表和LRU队列中删除  
  51.                         do_item_unlink_nolock(it, hv);  
  52.                         status = MOVE_DONE;  
  53.                     } else {  
  54.                         /* refcount == 1 + !ITEM_LINKED means the item is being 
  55.                          * uploaded to, or was just unlinked but hasn't been freed 
  56.                          * yet. Let it bleed off on its own and try again later */  
  57.                         status = MOVE_BUSY;  
  58.                     }  
  59.                 } else {//现在有worker线程正在引用这个item  
  60.                     status = MOVE_BUSY;  
  61.                 }  
  62.                 item_trylock_unlock(hold_lock);  
  63.             }  
  64.         }  
  65.   
  66.         switch (status) {  
  67.             case MOVE_DONE:  
  68.                 it->refcount = 0;//引用计数清零  
  69.                 it->it_flags = 0;//清零所有属性  
  70.                 it->slabs_clsid = 255;  
  71.                 break;  
  72.             case MOVE_BUSY:  
  73.                 refcount_decr(&it->refcount); //注意这里没有break  
  74.             case MOVE_LOCKED:  
  75.                 slab_rebal.busy_items++;  
  76.                 was_busy++;//记录是否有不能马上处理的item  
  77.                 break;  
  78.             case MOVE_PASS:  
  79.                 break;  
  80.         }  
  81.   
  82.         //处理这个页的下一个item  
  83.         slab_rebal.slab_pos = (char *)slab_rebal.slab_pos + s_cls->size;  
  84.         if (slab_rebal.slab_pos >= slab_rebal.slab_end)//遍历完了这个页  
  85.             break;  
  86.     }  
  87.   
  88.     //遍历完了这个页的所有item  
  89.     if (slab_rebal.slab_pos >= slab_rebal.slab_end) {  
  90.         /* Some items were busy, start again from the top */  
  91.         //在处理的时候,跳过了一些item(因为有worker线程在引用)  
  92.         if (slab_rebal.busy_items) {//此时需要从头再扫描一次这个页  
  93.             slab_rebal.slab_pos = slab_rebal.slab_start;  
  94.             slab_rebal.busy_items = 0;  
  95.         } else {  
  96.             slab_rebal.done++;//标志已经处理完这个页的所有item  
  97.         }  
  98.     }  
  99.   
  100.     pthread_mutex_unlock(&slabs_lock);  
  101.     pthread_mutex_unlock(&cache_lock);  
  102.   
  103.     return was_busy;//返回记录  
  104. }  

 

 

劫富济贫:

        上面代码中的was_busy就标志了是否有worker线程在引用内存页中的一个item。其实slab_rebalance_move函数的名字取得不好,因为实现的不是移动(迁移),而是把内存页中的item删除从哈希表和LRU队列中删除。如果处理完内存页的所有item,那么就会slab_rebal.done++,标志处理完成。在线程函数slab_rebalance_thread中,如果slab_rebal.done为真就会调用slab_rebalance_finish函数完成真正的内存页迁移操作,把一个内存页从一个slab class 转移到另外一个slab class中。

[cpp] view plain copy
 
 在CODE上查看代码片派生到我的代码片
    1. static void slab_rebalance_finish(void) {  
    2.     slabclass_t *s_cls;  
    3.     slabclass_t *d_cls;  
    4.   
    5.     pthread_mutex_lock(&cache_lock);  
    6.     pthread_mutex_lock(&slabs_lock);  
    7.   
    8.     s_cls = &slabclass[slab_rebal.s_clsid];  
    9.     d_cls   = &slabclass[slab_rebal.d_clsid];  
    10.   
    11.     /* At this point the stolen slab is completely clear */  
    12.     //相当于把指针赋NULL值  
    13.     s_cls->slab_list[s_cls->killing - 1] =  
    14.         s_cls->slab_list[s_cls->slabs - 1];  
    15.     s_cls->slabs--;//源slab class的内存页数减一  
    16.     s_cls->killing = 0;  
    17.   
    18.     //内存页所有字节清零,这个也很重要的  
    19.     memset(slab_rebal.slab_start, 0, (size_t)settings.item_size_max);  
    20.   
    21.     //将slab_rebal.slab_start指向的一个页内存馈赠给目标slab class  
    22.     //slab_rebal.slab_start指向的页是从源slab class中得到的。  
    23.     d_cls->slab_list[d_cls->slabs++] = slab_rebal.slab_start;  
    24.     //按照目标slab class的item尺寸进行划分这个页,并且将这个页的  
    25.     //内存并入到目标slab class的空闲item队列中  
    26.     split_slab_page_into_freelist(slab_rebal.slab_start,  
    27.         slab_rebal.d_clsid);  
    28.   
    29.     //清零  
    30.     slab_rebal.done       = 0;  
    31.     slab_rebal.s_clsid    = 0;  
    32.     slab_rebal.d_clsid    = 0;  
    33.     slab_rebal.slab_start = NULL;  
    34.     slab_rebal.slab_end   = NULL;  
    35.     slab_rebal.slab_pos   = NULL;  
    36.   
    37.     slab_rebalance_signal = 0;//rebalance线程完成工作后,再次进入休眠状态  
    38.   
    39.     pthread_mutex_unlock(&slabs_lock);  
    40.     pthread_mutex_unlock(&cache_lock);  
    41.   
    42. }  

posted on 2016-06-05 10:45  c++kuzhon  阅读(400)  评论(0)    收藏  举报

导航