awk分析日志

awk非常适合于文本文件处理,特别是利用内置数组,可以实现多个数据行关联处理。
比如,从应用日志中分析各类请求数量、处理时间等。

1,待处理日志文件a.trc
-------------------------------------------------------------------------
request=100001,amt=30,time=20110407152900
request=100002,amt=30,time=20110407152900
response=100001,time=20110407152910
request=100003,amt=30,time=20110407152900
response=100002,time=20110407152920
request=100004,amt=30,time=20110407152900
response=100003,time=20110407152930
request=100005,amt=30,time=20110407152900
response=100004,time=20110407152940
request=100006,amt=30,time=20110407152900
response=100005,time=20110407152950
request=100007,amt=30,time=20110407152900
response=100006,time=20110407152950
request=100008,amt=30,time=20110407152900
response=100007,time=20110407152950
-------------------------------------------------------------------------

3,处理结果:
-------------------------------------------------------------------------
reqid:100001 amt:30 bgn:20110407152900 end:20110407152910 ela:10 
reqid:100002 amt:30 bgn:20110407152900 end:20110407152920 ela:20 
reqid:100003 amt:30 bgn:20110407152900 end:20110407152930 ela:30 
reqid:100004 amt:30 bgn:20110407152900 end:20110407152940 ela:40 
reqid:100005 amt:30 bgn:20110407152900 end:20110407152950 ela:50 
reqid:100006 amt:30 bgn:20110407152900 end:20110407152950 ela:50 
reqid:100007 amt:30 bgn:20110407152900 end:20110407152950 ela:50 
resp:7 sumamt:240 
-------------------------------------------------------------------------

 

 

4.2 a2.awk
-------------------------------------------------------------------------
BEGIN {FS="=";
}
/request/ {reqp++;
           #print($0);
           vb=$2;sub(",amt","",vb);reqa[reqp,1]=vb;
           reqida[vb]=reqp;
           vb=$3;sub(",time","",vb);reqa[reqp,2]=vb;sumamt=sumamt+vb;
           reqa[reqp,3]=$4;
           next;
}

/response/{vb=$2;sub(",time","",vb);
           #printf("%s %s \n",vb,$0);
           vb3 = reqida[vb];
           if(vb3>0){reqa[vb3,4]=$3;}
           resp++;
           next;
}

END{
   for(i=1;i<=reqp;i++){
      t1 = reqa[i,3];
      t1f = substr(t1,1,4) " " substr(t1,5,2) " " substr(t1,7,2) " " substr(t1,9,2) " " substr(t1,11,2) " " substr(t1,13,2);
      t2 = reqa[i,4];
      t2f = substr(t2,1,4) " " substr(t2,5,2) " " substr(t2,7,2) " " substr(t2,9,2) " " substr(t2,11,2) " " substr(t2,13,2);
      if(length(t1)>0 && length(t2)>0) {ela = mktime(t2f)-mktime(t1f);} else {ela = 0;}
      printf("reqid:%s amt:%s bgn:%s end:%s ela:%d \n",reqa[i,1],reqa[i,2],reqa[i,3],reqa[i,4],ela);
   }
   printf("resp:%d sumamt:%d \n",resp,sumamt);
}

 

===================================================================================================================================================

next                  
Stop  processing  the current input record.  The next input record is read and processing starts over with the first pattern in the AWK program.  If the end  of  the input data is reached, the END block(s), if any, are executed.

举个例子:
cat file
1 a
2 b
3 c
4 d

awk '/^3/{print $2;next}{print $0}' file
1 a
2 b
c
4 d

如果匹配不到开头为3的记录,就打印$0
如果匹配到了开头为3的记录,就打印$2,这里如果没有next,会继续再打印$0

awk '/^3/{print $2}{print $0}'
1 a
2 b
c
3 c
4 d

next就是读取下一条记录,再从头执行代码

 

 

==========================================================================================================================================

AWK中内置变量的使用->用AWK每次处理多行文本

http://blog.csdn.net/imzoer/article/details/8738430

关于sed的多行模式空间,看这里

------------------------------------------------------------------------

AWK中,FS、NR、NF我们都熟悉了。那么RS是什么呢?

RS是记录分隔符。缺省为"\n"。

也就是说,awk是根据RS指定的符号作为一条记录的分隔符的。

在如下的文件中

zoer@ubuntu:~$ cat d
naughty 25 shandong;cc 24 guangdong;

我们以‘;’作为分隔符,那么使用awk打印第一个和第二个字段的代码:

[plain] view plaincopy
 
  1. zoer@ubuntu:~$ awk 'BEGIN{RS=";"}{print $1,$2}' d  
  2. naughty 25  
  3. cc 24  

在这里指定了RS为';'之后,可以正确打印了 。

那么OFS呢?OFS意思是 output field separator。也就是输出分隔符。在awk中,默认的输出分隔符是空格,我们可以改变默认的输出。比如说在上个例子中,我们把OFS改成换行。

[plain] view plaincopy
 
  1. zoer@ubuntu:~$ awk 'BEGIN{RS=";";OFS="\n"}{print $1,$2}' d  
  2. naughty  
  3. 25  
  4. cc  
  5. 24  

看得到,输出的内容与之前的例子的输出相比,已经换行了。

------------------------------------

ORS意思就是说,输出的记录分隔符了。例如我们把记录分隔符设置为####。

例子如下:

[plain] view plaincopy
 
  1. zoer@ubuntu:~$ awk 'BEGIN{RS=";";ORS="####"}{print $1,$2}' d  
  2. naughty 25####cc 24####  

 -------------------------------------------------------

awk默认每次处理一行文本。如果要每次处理多行文本,需要我们自己手动来实现。

下面,假设我们的数据一条记录是存在四行中的。并且每条记录之间用一个空行作为分隔。

要求输出第一个字段和最后一个字段。

看如下代码:

[plain] view plaincopy
 
  1. zoer@ubuntu:~$ cat data   
  2. naughty  
  3. 25  
  4. shandong linyi  
  5. naughty is a boy  
  6.   
  7. cc  
  8. 24  
  9. guangdong guangzhou  
  10. cc is a girl  
  11. zoer@ubuntu:~$ cat a  
  12. BEGIN {FS="\n";RS=""}  
  13. {print $1,$NF}  
  14. zoer@ubuntu:~$ awk -f a data  
  15. naughty naughty is a boy  
  16. cc cc is a girl  

从上面可以看到,我们设定了FS和RS,就可以做到了。RS=“”的意思是说将RS设置为一个空字符串。它代表了一个空行。

posted on 2015-10-03 16:41  小西红柿  阅读(582)  评论(0)    收藏  举报

导航