python模块

什么是模块?

  模块是一系列功能的集合体,

  来源:内置模块,第三方的模块,自定义模块

  模块的格式:

    1. 使用python编写的.py文件

    2. 已被编译为共享库或DLL的C或C++扩展

    3. 把一系列模块组织到一起的文件夹(注:文件夹下有一个__init__.py文件,该文件夹称之为包)

    4. 使用C编写并链接python解释器的内置模块

为什么要用模块?

  1. 提升开发效率

  2. 可以减少代码冗余

如何使用模块?
  前提:一定要区分谁是执行文件,谁是导入文件     

  import

  from ... import ...

1、import

  首次导入模块,会发生:

    1. 会产生一个模块的名称空间

    2. 执行文件,将执行中产生的名字放到模块的名称空间中

    3. 在当前执行文件的名称空间中拿到一个模块名,将名字指向模块的名称空间

  之后导入:都是直接引用第一次导入的成果,不会重新执行文件

  总结import导入模块:在使用时必须加上前缀:模块名.函数名

    优点:指名道姓地向某一个名称空间名字,肯定不会与当前名称空间中的名字冲突

    缺点:但凡应用模块中的名字都需要加前缀,不够简洁

# import dir.dir1导入了dir1和dir
import dir.dir1
dir.dir1.m2.f2()

dir.m1.f1()
示例

  注意:

    一行导入模块(不推荐)  # import spam, os, time

    可以为模块起别名(注意:模块名应该小写) import sdfafasasfasdfa as aaa

2、from ... import ...

  示例: from spam import money

  首次导入模块,会发生:

    1. 会产生一个模块的名称空间

    2. 执行文件,将执行中产生的名字放到模块的名称空间中

    3. 在当前执行文件拿到一个名字,该名字就是执行模块中相对应的 名字

  总结:

    优点:使用时无需加前缀,更简洁

    缺点:容易与当前名称空间名字冲突

  *代表所有,相当于取导入模块的__all__,通过设置在模块中设置__all__来控制*的内容

    from spam import * # __all__ = ['money', 'read1']   

#示范文件内容如下
#m1.py
print('正在导入m1')
from m2 import y
x='m1'


#m2.py
print('正在导入m2')
from m1 import x
y='m2'


#run.py
import m1


# 解决方法:
方法一:导入语句放到最后
#m1.py
print('正在导入m1')
x='m1'
from m2 import y


#m2.py
print('正在导入m2')
y='m2'
from m1 import x


方法二:导入语句放到函数中
#m1.py
print('正在导入m1')
def f1():
    from m2 import y
    print(x,y)
x = 'm1'
# f1()


#m2.py
print('正在导入m2')
def f2():
    from m1 import x
    print(x,y)
y = 'm2'


#run.py
import m1
m1.f1()
循环导入问题

  区分文件的两种用途:

    使用if __name__ == '__main__': pass  

# 扩展:
# 使用字符串导入

import importlib

s = 'demo.b'
res = importlib.import_module(s)  # 等价于from demo import b
# print(res.name)

# dir()内置函数, 查看当前名称空间内所有的变量
print(dir(res))

3、模块导入

  模块的搜索路径优先级:

    1. 内存中已经加载过的

    2. 内置模块

    3. sys.path,第一个值是,当前执行文件所在的文件夹 

  导入文件:环境变量以当前执行文件为准

    强调:所有被导入的模块参照环境变量sys.path都是以执行文件为准的

  绝对导入:

    以执行文件的sys.path为准导入,from 文件名1.文件名2... import 模块

    优点:执行文件与被导入模块都可以使用

    缺点:所有导入都是以sys.path为起点,导入麻烦

  相对导入:

    参照当前所在文件的文件夹为起始开始查找,称之为相对导入

    符号:. 表示当前文件,.. 表示上级文件, ... 表示上一级的上一级文件目录

    优点:导入更加简单

    缺点:只能在导入包中的模块时才能使用,不能执行文件中用

    强调:使用相对导入不能导入执行文件所在那一层的文件,不然抛出ValueError异常,

4、包

  包含__init__.py文件的文件夹,包是用来导入使用的,包内部包含的文件也都是导入使用的

  首次导入包,发生三件事:

    1. 以包下的__init__.py文件为基准来产生一个名称空间

    2. 执行包下的__init__.py文件的代码,将执行过程中产生的名字都丢到名称空间中

    3. 在当前执行文件中拿到一个名字p1,该p1指向__init__.py名称空间

  在py2中,包下必须有一个__init__.py文件,而python3即便没有也不会报错

  总结:

    1. 在导入语句中,. 的左边必须是一个包

    2. 导入包就是在导包下的__init__.py文件

    3. 如果使用绝对导入,绝对导入的起始位置都是以包的顶级目录为起始点

    4. 但是包内部模块的导入通常应该使用相对导入,用 . 代表当前文件(而非执行文件),

     这样如果要修改顶级目录名,则包内的就不用做任何修改

    强调:

      1. 相对导入只能在包内部的模块之间相互导入

      2. .. 上一级不能超出包的顶级目录

5、常用模块

  1. time 模块

    (python调用系统的时间模块)

    在python中时间分为3种:

      1. 时间戳, timestamp 从1970年1月1日到现在的秒数,主要用于计算

      2. localtime,本地时间,表示是计算机当前所在位置的时间

      3. UTC 世界协调时间

    常用方法:

      import time

      time.time(), 时间戳

      time.localtime(), 获取当地时间,返回的是结构化时间,或把时间戳转成结构化时间。

      time.gmtime(), 获取UTC时间,也是结构化时间,比中国时间少8个小时

      time.strptime(), 将格式化字符串的时间转为结构化时间

      time.strftime(format, t), format 是期望的格式,只支持结构化时间和元组格式

      time.mktime(), 结构化转成时间戳

      time.sleep(), 线程推迟指定时间后运行,单位是秒

    不常用的方法:

      time.asctime(t), 把一个元组或结构化时间变成美国时间格式

      time.ctime(sec), 把一个时间戳变成美国时间格式

import time
# 获取时间戳  返回浮点型
print(time.time())
# 结果是:1553428239.7184734

# 获取当地时间   返回的是结构化时间
print(time.localtime())
# 结果是:time.struct_time(tm_year=2019, tm_mon=3, tm_mday=24, tm_hour=19,
# tm_min=18, tm_sec=58, tm_wday=6, tm_yday=83, tm_isdst=0)

#  获取UTC时间 返回的还是结构化时间  比中国时间少8小时
print(time.gmtime())

# 将获取的时间转成我们期望的格式 仅支持结构化时间或时间格式的元组
print(time.strftime("%Y-%m-%d %H:%M:%S",time.localtime()))
# 结果是:'2019-03-24 19:42:13'
>>>time.strftime("%Y-%m-%d %H:%M:%S", (2019,4,6,11,17,5,0,0,0))
'2019-04-06 11:17:05'


# 将格式化字符串的时间转为结构化时间  注意 格式必须匹配
print(time.strptime("2018-08-09 09:01:22","%Y-%m-%d %H:%M:%S"))
# 结果是:time.struct_time(tm_year=2018, tm_mon=8, tm_mday=9, tm_hour=9, 
# tm_min=1, tm_sec=22, tm_wday=3, tm_yday=221, tm_isdst=-1)

# 时间戳 转结构化
print(time.localtime(time.time()))

# 结构化转 时间戳
print(time.mktime(time.localtime()))

# sleep 让当前进程睡眠一段时间 单位是秒
# time.sleep(2)
# print("over")


# 不太常用的时间格式
# print(time.asctime())
# print(time.ctime())
# 结果是:'Sun Mar 24 19:45:26 2019'
使用示例
%a        weekday缩写
%A        weekday全写
%b        mouth name 缩写  
%B        mouth name 全写  
%c        美国时间格式
%d        日期[0,31]
%M        分
%m        月(数字)
%I        (12小时制)的小时
%H        (24小时制)的小时
%j        一年的第多少天
%p        AM或PM
%s        秒
%u        第多少个星期
%w        星期(数字)
%x        日期,格式如:“03/24/19”
%y        年数,如:19
%Y        年数,如:2019
%z        加上“+0800”
字符串转时间格式对应表

   2. datetime模块

    (python内置的时间模块)

    datetime.datetime.now(), 返回当前时间,格式化时间对象,

    datetime.datetime.now().timestamp(),获得当前时间戳

    datetime.datetime.fromtimestamp(), 可以从时间戳转换成datetime类型

    datatime.timedelta(), 表示时间差,最多只能到天

    两个时间差可以+-*/,时间差和datetime可以 +-

import datetime

# 获取时间 获取当前时间 并且返回的是格式化字符时间
print(datetime.datetime.now())

# 单独获取某个时间 年 月
d = datetime.datetime.now()
print(d.year)
print(d.day)

# 手动指定时间
d2 = datetime.datetime(2018,8,9,9,50,00)
print(d2)

# 计算两个时间的差  只能- 不能加+
print(d - d2)

# 替换某个时间
print(d.replace(year=2020))

# 表示时间差的模块 timedelta
print(datetime.timedelta(days=1))

t1 = datetime.timedelta(days=1)
t2 = datetime.timedelta(weeks=1)
print(t2 - t1)
# 时间差可以和一个datetime进行加减
print(d + t2)
使用示例

  3. random模块

    random.random(),  产生一个0到1的额随机浮点数,区间 [0,1) 

    random.uniform(1,10),  产生一个1到10的随机浮点数,区间 [1,10)

    random.randint(1,100),  产生一个1到100的随机数,区间[1,100]

    random.randrange(1,100),  产生一个1到100的随机数,区间[1,100), 可以加步长

      如:random.randrange(1, 100, 2),在1~100之间产生一个随机奇数

    random.choices('abc123@#$'), 从给定的数据集合中返回一个随机数,返回列表 ["c"]

    random.sample('abcdefg',3)从给定的字符中选取3个字符,返回列表 ["d", "f", "g"]

    random.shuffle(),  把数据顺序打乱

# 快速产生一个随机数
# 使用了string模块
# string.digits, '0123456789'
# string.ascii_letters, 'a~zA~Z'
# string.ascii_lowercase, 'a~z'
# string.ascii_uppercase, 'A~Z'
# string.punctuation, 特殊字符
# string.printable, 可打印的所有字符,包括数字大小写字母,特殊字符

# 产生四位随机字母或数字的字符串
import random
import string
''.join(random.sample(string.digits+string.ascii_letters, 4))
快速产生一个四位的随机字母或数字的字符串

  4. sys模块

    系统相关的模块,一般用于脚本程序

    sys.argv 获取cmd命令传进的参数,第一个参数是程序本身的路径

    sys.exit(n), 退出程序,正常退出时exit(0)

    sys.version, 获取python解释器的版本信息

    sys.platform,返回操作系统平台名称,window返回'win32', centos7返回'linux2'

# 自定义进度条
def process_bar(percent, width=30, msg='进度: '):
    """
    自定义进度条
    格式:"进度: [******************************]100%"
    :param percent: 百分比[0,1]
    :param width: 打印宽度,默认:30
    :param msg: 进度条提示,默认:"进度: "
    :return:
    """
    percent = percent if percent<=1 else 1
    prompt = ("\r%s[%%-%ds]"%(msg, width))%('*' * int(width*percent))
    prompt += "%s%%"
    print(prompt%(int(percent*100)), end='')
自定义进度条

  5. shutil模块

    shutil.copyfileobj(fsrc, fdst, length), 拷贝文件,提供两个文件对象,长度代表缓存区

    shutil.copyfile(src, dst), 拷贝文件,提供两个文件路径

    shutil.copymode(),拷贝文件权限,提供两个文件路径

    shutil.copystat(src,dst), 拷贝文件状态信息,包含mode, bits, atime, mtime, flags

    shutil.copy(src, dst) 拷贝文件和权限,提供两个文件路径

    shutil.copy2(src, dst) 拷贝文件和状态信息,提供两个文件路径

    shutil.ignore_patterns(*patterns)

    shutil.copytree(src, dst,symlinks=False, ignore=None) 拷贝目录,

      symlinks: 在linux下,默认为False将软连接拷贝为硬链接,否则拷贝成软连接

      ignore=shutil.ignore_patterns("mp3", '.py')

    shutil.rmtree(path, ignore_error) 删除目录 可以设置忽略文件,递归删除

    shutil.move(src, dst),移动目录和文件

    压缩和解压(shutil模块只能压缩,不能解压)

      shutil.make_archive(base_name, format, root_dir,...)

        base_name: 压缩包的名字或路径

        format: 压缩包的种类,'zip', 'tar', 'bztar', 'gztar'

        root_dir: 要压缩的文件路径

import shutil
import zipfile
import tarfile

# 利用shutil来创建压缩文件  仅支持 tar 和zip格式 内部调用zipFIle tarFIle模块实现
# shutil.make_archive("test","zip",root_dir="D:\aaa")

# 解压zip
# z = zipfile.ZipFile(r"D:\test.zip")
# z.extractall()
# z.close()

# 解压tar
# t = tarfile.open(r"D:\test.tar")
# t.extractall()
# t.close()

# 使用zip压缩
# z = zipfile.ZipFile("myzip.zip", 'w')
# z.write('a.py')
# z.write('date') # 如果是目录,只能写目录,目录下的内容未写进去

# 使用tar压缩
# t = tarfile.open(r"D:\mytar.tar","w")
# t.add(r"D:\datetime_test.py")
# t.add(r"D:\time_test.py")
# t.close()
压缩和解压(shutil模块,zipfile模块,tarfile模块)

  6. os模块

    与操作系统相关的一些操作

    os.getcwd(),获取当前工作目录,即解释器所在目录=os.path.dirname(__file__)

    os.chdir('dirname'), 改变当前脚本工作目录,

    os.curdir,返回当前目录('.')

    os.pardir, 获取当前目录的父目录字符串名('..')

    os.makedirs('dir1/dir2'), 递归创建目录

    os.remakedirs('dir1/dir2'), 递归删除目录,若目录为空,则删除,并递归到上一级目录

    os.mkdir('dirname'), 生成单极目录

    os.rmdir('dirname'), 删除单极目录,若目录不为空,则无法删除,报错

    os.listdir('dirname'), 列出指定目录下的所有文件和子目录,包括隐藏文件

    os.remove() 删除一个文件

    os.rename("oldname", "newname")

    os.stat('path/filename'), 获取文件/目录信息

    os.sep, 获取操作系统指定的路径分隔符,win为'\\', linux为'\'

    os.linesep, 获取操作系统指定的 行终止符, win为'\r\n', linux为'\n'

    os.pathsep, 获取用分割文件路径的字符串,win下为';', linux下为':'

    os.name, 获取当前平台,win为'nt', linux为'posix'

    os.system('bash command '), 运行shell命令

    os.environ 获取系统环境变量

    os下的path模块

      from os import path

      path.abspath(path), 获取文件的绝对路径

      path.split(path), 将path分割成目录和文件名二元组返回

      path.splitext(), 分离后缀

      path.dirname(path),返回path的目录,即path.split(path)的第一个元素

      path.exists(),path存在,就返回True

      path.isabs(path), 判断是否是绝对路径

      path.isdir(path), 判断是否是目录

      path.isfile(path), 判断是否是文件

      path.join(path1,path2,..), 将多个路径组合后返回

      path.getatime(path), 返回文件或目录最后存取时间

      path.getmtime(path), 返回文件或目录最后修改时间

      path.getsize(path),返回 文件或目录大小

      path.normcase(), 用于规范化文件(包括路径),将大写转换成小写, 斜杆变成当前系统分隔符

      path.normpath(), 用于规范化路径,将斜杆变成当前系统分隔符,可以识别'.'和'..'

  7. pickle模块

    用于序列化,把内存中的数据变成字符串

    序列化

      pickle.dumps(d)

      pickle.dump(d, f),f为文件对象

    反序列化

      pickle.loads(d)

      pickle.load(f)

  8. json模块

    能存储的有str, int, float, dic, list, bool

    json保存的字符串中的引号必须是双引号

    最好不要多次dump,把数据都放到一起在dump

    方法:

      json.dumps(d)

      json.dump(d, f, indent=4), indent表示写入文件时加上格式

      json,load(d)

      json.loads(f)

    json类型          python类型
    {}(对象)               dict
    [](数组)            list, tuple
     string                  'str'
     123.45             int或float
    true/false         True/False
      null                  None

>>> a = json.dumps({'a': True,'s': (1,2,4)})
>>> a
'{"a": true, "s": [1, 2, 4]}'
>>> b= json.loads(a)
>>> b
{'a': True, 's': [1, 2, 4]}        
json类型和python数据类型对应关系
import json
from datetime import datetime,date

class CustomJson(json.JSONEncoder):
    def default(self, o):
        if isinstance(0, datetime):
            return o.strftime("%Y-%m-%d %X")
        elif isinstance(o, date):
            return o.strftime("%Y-%m-%d")
        else:
            super().default(self, o)

res = {'c1': datetime.today(), 'c2': date.today()}
print(json.dumps(res, cls=CustomJson))
使用json内置的类自定数据类型转换

  9. shelve模块

    只有一个open函数,返回类型字典的对象,可以多次写入读出

import shelve
# 写数据,得到了三个文件后缀分别是.bak, .dat, .dir
# f = shelve.open(r'test_shelve')
# f['stu1'] = {'name': 'aaa', 'age': 22, 'gender': "boy"}
# f['stu2'] = {'name': 'bbb', 'age': 12, 'gender': "boy"}
# f.close()

# 读数据
# f = shelve.open(r'test_shelve')
# print(f['stu1'])
# for key in f:
#     print(key, f[key])
# f.close()

# 可以再次写入,或修改
f = shelve.open(r'test_shelve')
f['stu3'] = {'name': 'ccc', 'age': 12, 'gender': "girl"}
for key in f:
    print(key, f.get(key))
f.close()
使用示例

  10. xml模块

<?xml version="1.0"?>
<data>
    <country name="Liechtenstein">
        <rank updated="yes">2</rank>
        <year>2008</year>
        <gdppc>141100</gdppc>
        <neighbor name="Austria" direction="E"/>
        <neighbor name="Switzerland" direction="W"/>
    </country>
    <country name="Panama">
        <rank updated="yes">69</rank>
        <year>2011</year>
        <gdppc>13600</gdppc>
        <neighbor name="Costa Rica" direction="W"/>
        <neighbor name="Colombia" direction="E"/>
    </country>
</data>
XML的格式
# 查找:
import xml.etree.ElementTree as ET
# 打开文件, 类似open
tree = ET.parse('test_xml')
# f,seek(0),得到的root是<Element 'data' at 0x000002106B266818>
root = tree.getroot()
print(root.tag) # 结果是 data

# 遍历xml文档
for child in root:
    print('========>', child.tag, child.attrib, child.attrib['name'])
    for i in child:
        print(i.tag, i.attrib, i.text)

# 只遍历year节点
for node in root.iter('year'):
    print(node.tag,node.text)

# 结果:
# ========> country {'name': 'Liechtenstein'} Liechtenstein
# rank {'updated': 'yes'} 2
# year {} 2008
# gdppc {} 141100
# neighbor {'name': 'Austria', 'direction': 'E'} None
# neighbor {'name': 'Switzerland', 'direction': 'W'} None
# ========> country {'name': 'Panama'} Panama
# rank {'updated': 'yes'} 69
# year {} 2011
# gdppc {} 13600
# neighbor {'name': 'Costa Rica', 'direction': 'W'} None
# neighbor {'name': 'Colombia', 'direction': 'E'} None

# year 2008
# year 2011
查找xml的标签和属性
# 修改
# import xml.etree.ElementTree as ET
# tree = ET.parse('test_xml')
# root = tree.getroot()
#
# 修改text和attrib
# for node in root.iter('year'):
#     new_year = int(node.text) + 1
#     node.text = str(new_year) # 必须是字符串
#     node.set('updated', 'yes')# 添加新属性,可以是数字
#     node.set('version', '2.0')
# tree.write('test_xml1') # 最好写入到一个新的文件内

# 修改标签tag
# for node in root.iter('gdppc'):
#     node.tag = node.tag.upper()
# tree.write('test_xml1')

# 删除node, findall 找到所有的节点,find只找第一个,就结束
# for country in root.findall('country'):
#     rank = int(country.find('rank').text)
#     if rank > 50:
#         root.remove(country)
# tree.write('test_xml1')
修改xml的标签和属性
# 插入
import xml.etree.ElementTree as ET
tree = ET.parse("test_xml")
root=tree.getroot()
for country in root.findall('country'):
    for year in country.findall('year'):
        if int(year.text) > 2000:
            year2=ET.Element('year2')
            year2.text='新年'
            year2.attrib={'update':'yes'}
            country.append(year2) #往country节点下添加子节点
tree.write('test_xml')
再xml内插入新的标签
# 创建
import xml.etree.ElementTree as ET

new_xml = ET.Element("namelist")
name = ET.SubElement(new_xml, "name", attrib={"enrolled": "yes"})
age = ET.SubElement(name, "age", attrib={"checked": "no"})
sex = ET.SubElement(name, "sex")
sex.text = '33'
name2 = ET.SubElement(new_xml, "name", attrib={"enrolled": "no"})
age = ET.SubElement(name2, "age")
age.text = '19'

et = ET.ElementTree(new_xml)  # 生成文档对象
et.write("test_xml2.xml", encoding="utf-8", xml_declaration=True)
# xml_declaration:版本号说明

ET.dump(new_xml)  # 打印生成的格式
创建一个新的xml文件

  11. configparser模块

# 注释1
; 注释2

[section1]
k1 = v1
k2:v2
user=egon
age=18
is_admin=true
salary=31

[section2]
k1 = v1
配置文件格式后缀为ini
# 查找
import configparser
config=configparser.ConfigParser()
config.read('a.cfg')

#查看所有的标题
res=config.sections()
print(res) #['section1', 'section2']

#查看标题section1下所有key=value的key
options=config.options('section1')
print(options) #['k1', 'k2', 'user', 'age', 'is_admin', 'salary']

#查看标题section1下所有key=value的(key,value)格式
item_list=config.items('section1')
print(item_list) #[('k1', 'v1'), ('k2', 'v2'), ('user', 'egon'), ('age', '18'), ('is_admin', 'true'), ('salary', '31')]

#查看标题section1下user的值=>字符串格式
val=config.get('section1','user')
print(val) #egon

#查看标题section1下age的值=>整数格式
val1=config.getint('section1','age')
print(val1) #18

#查看标题section1下is_admin的值=>布尔值格式
val2=config.getboolean('section1','is_admin')
print(val2) #True

#查看标题section1下salary的值=>浮点型格式
val3=config.getfloat('section1','salary')
print(val3) #31.0
查找属性
# 修改
import configparser
config=configparser.ConfigParser()
config.read('a.cfg',encoding='utf-8')

#删除整个标题section2
config.remove_section('section2')

#删除标题section1下的某个k1和k2
config.remove_option('section1','k1')
config.remove_option('section1','k2')

#判断是否存在某个标题
print(config.has_section('section1'))

#判断标题section1下是否有user
print(config.has_option('section1',''))

#添加一个标题
config.add_section('sun')

#在标题sun下添加name=sun,age=18的配置
config.set('sun','name','sun')
config.set('sun','age',18) #报错,必须是字符串

#最后将修改的内容写入文件,完成最终的修改
config.write(open('a.cfg','w'))
修改文件内容
# 创建一个ini文件
# 读取到的ini文件类似字典的格式,所以可以以字典的方式写入
import configparser
  
config = configparser.ConfigParser()
config["DEFAULT"] = {'ServerAliveInterval': '45',
                      'Compression': 'yes',
                     'CompressionLevel': '9'}
  
config['bitbucket.org'] = {}
config['bitbucket.org']['User'] = 'hg'
config['topsecret.server.com'] = {}
topsecret = config['topsecret.server.com']
topsecret['Host Port'] = '50022'     # mutates the parser
topsecret['ForwardX11'] = 'no'  # same here
config['DEFAULT']['ForwardX11'] = 'yes'
with open('example.ini', 'w') as configfile:
   config.write(configfile)
创建一个新的ini文件

  12. hashlib模块

  hash:一种算法,主要有SHA1, SHA224, SHA256, SHA384, SHA512, MD5算法

  功能:

    1. 输入任意长度的信息,输出一定位数的随机字符串(数字指纹)

    2. 不同的输入得到不同的结果(唯一性)

  特点:

    1. 压缩性:任意长度的数据,一种算法的得到的结果长度是固定的

    2. 抗修改性:对原数据的修改,新生成的hash值区别也会很大

    3. 强抗碰撞:已知原数据和hash值,想找到一个具有相同hash值的数据是非常困难的

    4. 不可逆性:不能通过hash值得到原数据

import hashlib
m = hashlib.md5()
m.update(b"hello")    
print(m.hexdigest())  # #5d41402abc4b2a76b9719d911017c592

m.update('alvin'.encode('utf8'))
print(m.hexdigest())  #6b2ca1ef751357e9bf4ad0b52b6dead9

# 注意:把一段很长的数据update多次,与一次update这段长数据,得到的结果一样
# update多次为校验大文件提供了可能。
MD5的使用
# 另一种hash方法

# import hmac
# h = hmac.new('sun'.encode('utf8'))
# h.update('hello'.encode('utf8'))
# print (h.hexdigest())#d62c8d48809fff8382a346d9bd29f942

#要想保证hmac最终结果一致,必须保证:
#1:hmac.new括号内指定的初始key一样
#2:无论update多少次,校验的内容累加到一起是一样的内容

import hmac
h1=hmac.new(b'sun')
h1.update(b'abc')
h1.update(b'123')
print(h1.hexdigest())


h2=hmac.new(b'sun')
h2.update(b'abc123')
print(h2.hexdigest())


h3=hmac.new(b'sunabc123')
print(h3.hexdigest())


# 结果: 
# c6db20c8a208f27262c392618889ae14
# c6db20c8a208f27262c392618889ae14
# 2c5236ca81f8a1cf527834430de24b07
另一种hash模块,hamc使用了key-value方式

  13. subprocess模块

    作用:用于执行系统命令

    常用方法:(都是对Popen方法的封装)

      run   返回一个表示执行结果的对象

      call   返回的执行的状态码

import subprocess
# print(1)
# res = subprocess.run("tasklist",shell=True,stdout=subprocess.PIPE)
#
# print(res.stdout.decode("gbk"))
#
# print(res.stderr)

# res = subprocess.call("tasklist",shell=True)
# print(res)

#  第一个进程a读取tasklist的内容   将数据交给另一个进程b  进程b将数据写到文件中
res1 = subprocess.Popen("tasklist",stdout=subprocess.PIPE,shell=True,stderr=subprocess.PIPE)
# print("hello")
#print(res1.stdout.read().decode("gbk"))
#print(res1.stderr.read().decode("gbk"))
#  得到的结果一旦读出来就不能再读
res2 = subprocess.Popen("findstr cmd",stdout=subprocess.PIPE,shell=True,stderr=subprocess.PIPE,stdin=res1.stdout)
print(res2.stdout.read().decode("gbk"))
使用示例

  14. logging模块

    日志级别

      logging.debug("debug日志")          # 代表的数字:10

      logging.info("info日志")                  # 20

      logging.warning("warning日志")    # 30

      logging.error("error日志")              # 40

      logging.critical("critical日志")         # 50

    默认级别:warning

# 基本使用
import logging
logging.basicConfig(filename='access.log',
                    format='%(asctime)s - %(name)s - %(levelname)s -%(module)s:  %(message)s',
                    datefmt='%Y-%m-%d %H:%M:%S %p',
                    level=10)

logging.debug('调试debug')
logging.info('消息info')
logging.warning('警告warn')
logging.error('错误error')
logging.critical('严重critical')

#========结果
access.log内容:
2017-07-28 20:32:17 PM - root - DEBUG -test:  调试debug
2017-07-28 20:32:17 PM - root - INFO -test:  消息info
2017-07-28 20:32:17 PM - root - WARNING -test:  警告warn
2017-07-28 20:32:17 PM - root - ERROR -test:  错误error
2017-07-28 20:32:17 PM - root - CRITICAL -test:  严重critical
#===========问题
# 1. 不能修改写入文件的编码方式
# 2. 不能打印到终端

# basicConfig内参数的含义:
filename:用指定的文件名创建FiledHandler,这样日志会被存储在指定的文件中。
filemode:文件打开方式,在指定了filename时使用这个参数,默认值为“a”还可指定为“w”。
format:指定handler使用的日志显示格式。
datefmt:指定日期时间格式。
level:设置rootlogger的日志级别
stream:为Ture时,打印到终端;若同时列出了filename和stream两个参数,则stream参数会被忽略。
简单使用
%(name)s Logger的名字
%(levelno)s 数字形式的日志级别
%(levelname)s 文本形式的日志级别
%(pathname)s 调用日志输出函数的模块的完整路径名,可能没有
%(filename)s 调用日志输出函数的模块的文件名
%(module)s 调用日志输出函数的模块名
%(funcName)s 调用日志输出函数的函数名
%(lineno)d 调用日志输出函数的语句所在的代码行
%(created)f 当前时间,用UNIX标准的表示时间的浮 点数表示
%(relativeCreated)d 输出日志信息时的,自Logger创建以 来的毫秒数
%(asctime)s 字符串形式的当前时间。默认格式是 “2003-07-08 16:49:45,896”。逗号后面的是毫秒
%(thread)d 线程ID。可能没有
%(threadName)s 线程名。可能没有
%(process)d 进程ID。可能没有
%(message)s用户输出的消息      
format参数的格式化类型

    日志下的四个概念(就是对象):

      logger对象:负责产生各种级别的日志,提供应用程序可以直接使用的接口

       filter对象:对日志经行筛选

      handler对象:分发日志, 控制日志输出目标位置

        formatter对象:决定日志记录的最终输出格式

import logging
import time
# 1. logger对象:负责产生日志,然后交给Filter过滤,然后交给不同的Handler输出
logger=logging.getLogger("登陆日志")

# 2. Filter对象:不常用,略

# 3. Handler对象:控制日志输出的目标位置,可以更改文件的写的模式和编码的模式
h1 = logging.FileHandler('t1.log', encoding='utf-8')
h2 = logging.FileHandler('t2.log', encoding='utf-8')
h3 = logging.StreamHandler()

# 4. Formatter对象,日志格式
fmtter1 = logging.Formatter('%(asctime)s - %(name)s - %(levelname)s -%(module)s:  %(message)s',
                            datefmt='%Y-%m-%d %H:%M:%S %p')
fmtter2 = logging.Formatter('%(asctime)s - %(name)s:  %(message)s',
                            datefmt='%Y-%m-%d %H:%M:%S %p')

# 5. 为Handler对象绑定格式
h1.setFormatter(fmtter1)
h2.setFormatter(fmtter2)
h3.setFormatter(fmtter2)

# 6. 将handler添加给logger
logger.addHandler(h1)
logger.addHandler(h2)
logger.addHandler(h3)

# 7. 设置日志级别,有logger对象与handler对象两层关卡
#    必须都放行最终日志才会放行,通常两者级别相同
logger.setLevel(logging.DEBUG)
h1.setLevel(10)
h2.setLevel(10)
h3.setLevel(10)

# 8. 使用logger对象产生日志
logger.info("sun登陆了终端")
time.sleep(3)
logger.critical("sun输入了 rm -rf *")
logger.info("sun注销了用户,并退出")
logger.critical("sun 跑路了。。。")
自定义日志输出
# 第二种方法使用配置文件
# settings.py文件内

# Formatter下的日志格式
standard_format = '%(asctime)s - task:%(name)s - %(filename)s:%(lineno)d -' \
                  ' %(levelname)s : [%(message)s]'
simple_format = '%(filename)s:%(lineno)d - %(levelname)s : [%(message)s]'

# 写入文件时,文件的绝对路径
fh1_path = r'a1.log'
fh2_path = r'a2.log'

# log配置字典
LOGGING_DIC = {
    'version': 1,
    'formatters': {
        # standard,simple为日志格式名字,可任意
        'standard': {
            'format': standard_format
        },
        'simple': {
            'format': simple_format
        },
    },
    'filters': {},
    'handlers': {
        #打印到终端的日志
        'ch': {
            'level': 'DEBUG',
            'class': 'logging.StreamHandler',  # 打印到终端
            'formatter': 'simple'
        },
        #打印到a1.log文件的日志
        'fh1': {
            'level': 'DEBUG',
            'class': 'logging..handlers.RotatingFileHandler',  # 保存到文件,RotatingFileHandler可以设置最大存储容量maxBytes
            'formatter': 'standard',
            'filename': fh1_path,  # 日志文件的路径
            'maxBytes': 1024*1024*5,  # 日志大小 5M
            'backupCount': 5,        # 文件个数
            'encoding': 'utf-8',  # 日志文件的编码,再也不用担心中文log乱码了
        },
        # 打印到a2.log文件的日志
        'fh2': {
            'level': 'DEBUG',
            'class': 'logging.FileHandler',  # 保存到文件
            'formatter': 'simple',
            'filename': fh2_path,  # 日志文件的路径
            'maxBytes': 1024*1024*5,  # 日志大小 5M
            'backupCount': 5,        # 文件个数
            'encoding': 'utf-8',  # 日志文件的编码,再也不用担心中文log乱码了
        },
    },
    'loggers': {
        # key为空字符串,意味着如果找不到日志名,就以空字符串的格式为准
        '': {
            'handlers': ['fh1', 'fh2', 'ch'],
            'level': 'DEBUG',
        },
    },
}

# run.py下运行
import logging.config
import settings

logging.config.dictConfig(settings.LOGGING_DIC)

logger1=logging.getLogger('用户交易')
#logger1-> fh1,fh2,ch
logger1.info('要钱没有,要命一条')

logger2=logging.getLogger('用户权限')
#logger2-> fh1,fh2,ch
logger2.error('sun查看了自己一个亿的账号')
把日志格式放到配置文件内
# 写项目的小技巧
    把项目文件夹放到根目录下, 如:D:/,且把执行文件放到顶级目录下。
    目的:方便查找,不用再把文件路径添加到sys.path内,使用pycharm导入模块时也可以有提示
# 使用logging技巧
    把配置信息放到conf下的settings.py,把生成日志的内容放到lib下的common.py文件  内,这样每次要用日志时直接调用即可
写项目的小技巧

  15. re模块

     就是一些带有特殊含义的符号或符号的组合,内部实现 不是python 而是调用了c库

    作用:对字符串进行过滤

元字符
描述
. 
匹配任意字符,除了换行符,当re.DOTALL标记被指定时,则可以匹配包括换行符的任意字符
^
匹配字符串开头,遇到换行结束,当flag=re.MULTILINE,可以在多行中查找
$
匹配字符串结尾,遇到换行结束,当flag=re.MULTILINE,可以在多行中查找
*
匹配*号前的字符0或多次
+
匹配+号前的字符1或多次
?
匹配0或1个由前面的正则表达式定义的片段,非贪婪方式
{n}
匹配n个前面的表达式
{,n}
匹配0到n个前面的表达式
{n,m}
匹配n到m次由前面的正则表达式的片段,先取m次,没有再依次递减,贪婪模式
|
匹配|左或|右的字符
[...]
用来表示从一组字符取出一个
[^...]
不再[]内的字符
( ) 
匹配括号内的表达式,也表示一个组,不会改变原来的表达式逻辑意义
()就是提高优先级,可使用(?:)取消优先级
\
转义字符
\A
匹配字符串开头,和^用法一样
\Z
匹配字符串结尾,和$用法一样
\d
匹配任意数字,等价于[0-9]
\D 
匹配非任意数字,等价于[^0-9]
\w
匹配字母数字下划线,等价于[_A-Za-z0-9]
\W
匹配非字母数字下划线,等价于[^_A-Za-z0-9]
\s
匹配任意空字符,等价于[\t\n\r]
\S
匹配任意非空字符,等价于[^\t\n\r]
\b
匹配一个单词边界,也就是单词的末尾
\B
匹配一个非单词边界
(?<name>..)
分组匹配, 可以使用字典格式取出数字

 

示例:
import re
# ^与$
print(re.findall('^h','hello egon 123')) #['h']
print(re.findall('3$','hello egon 123')) #['3']

# 重复匹配:| . | * | ? | .* | .*? | + | {n,m} |
# .
print(re.findall('a.b','a1b a*b a b aaab')) #['a1b', 'a*b', 'a b', 'aab']
print(re.findall('a.b','a\nb')) #[]
print(re.findall('a.b','a\nb',re.S)) #['a\nb'],
print(re.findall('a.b','a\nb',re.DOTALL)) #['a\nb']同上一条意思一样
# *
print(re.findall('ab*','bbbbbbb')) #[]
print(re.findall('ab*','a')) #['a']
print(re.findall('ab*','abbbb')) #['abbbb']
# ?
print(re.findall('ab?','a')) #['a']
print(re.findall('ab?','abbb')) #['ab']

#匹配所有包含小数在内的数字
print(re.findall('\d+\.?\d*',"asdfasdf123as1.13dfa12adsf1asdf3")) #['123', '1.13', '12', '1', '3']

# .*默认为贪婪匹配, 会一直匹配到不满足条件为止 ,用?来阻止
print(re.findall('a.*b','a1b22222222b')) #['a1b22222222b']

#.*?为非贪婪匹配:推荐使用,找最小的
print(re.findall('a.*?b','a1b22222222b')) #['a1b']

# +
print(re.findall('ab+','a')) #[]
print(re.findall('ab+','abbb')) #['abbb']

# {n,m}
print(re.findall('ab{2}','abbb')) #['abb']
print(re.findall('ab{2,4}','abbb')) #['abb']
print(re.findall('ab{1,}','abbb')) #'ab{1,}' ===> 'ab+'
print(re.findall('ab{0,}','abbb')) #'ab{0,}' ===> 'ab*'
print(re.findall('ab{,2}','abbb')) #'ab{,2}' ===> 'ab{0,2}'

# []
print(re.findall('a[1*-]b','a1b a*b a-b')) #[]内的都为普通字符了,且如果-没有被转意的话,应该放到[]的开头或结尾, 结果['a1b','a*b','a-b']
print(re.findall('a[^1*-]b','a1b a*b a-b a=b')) #[]内的^代表的意思是取反,以结果为['a=b']
print(re.findall('a[0-9]b','a1b a*b a-b a=b')) #[]内表示0到9任意数,所以结果为['a1b']
print(re.findall('a[a-z]b','a1b a*b a-b a=b aeb')) #[]内的^代表的意思是取反,所以结果为['aeb']
print(re.findall('a[^a-zA-Z]b','a1b a*b a-b a=b aeb aEb')) #['a1b','a*b','a-b','a=b']

# ()
print(re.findall('(ab)+123','ababab123')) #['ab'],匹配到末尾的ab123中的ab
print(re.findall('(?:ab)+123','ababab123')) #['ababab123']findall的结果不是匹配的全部内容,而是组内的内容,?:可以让结果为匹配的全部内容

print(re.findall('href="(.*?)"','<a href="http://www.baidu.com">点击</a>'))#['http://www.baidu.com']
print(re.findall('href="(?:.*?)"','<a href="http://www.baidu.com">点击</a>'))#['href="http://www.baidu.com"']

# |
print(re.findall('compan(?:y|ies)','Too many companies have gone bankrupt, and the next one is my company')) #结果:['companies', 'company']

# \
print(re.findall(r'a\\c','a\c')) #r代表告诉解释器使用rawstring,即原生字符串,把我们正则内的所有符号都当普通字符处理,不要转义
print(re.findall('a\\\\c','a\c')) #同上面的意思一样,和上面的结果一样都是['a\\c']

# \w与\W
print(re.findall('\w','sun 123')) #['s', 'u', 'n', '1', '2', '3']
print(re.findall('\W','sun 123')) #[' ']

#\s与\S
print(re.findall('\s','sun 123')) #[' ']
print(re.findall('\S','sun 123')) #['s', 'u', 'n', '1', '2', '3']

#\n \t都是空,都可以被\s匹配
print(re.findall('\s','hello \nsun \t123')) #[' ', '\n', ' ', '\t']

#\n与\t
print(re.findall(r'\n','sun \n123')) #['\n']
print(re.findall(r'\t','sun\t123')) #['\n']

#\d与\D
print(re.findall('\d','sun 123')) #['1', '2', '3']
print(re.findall('\D','sun 123')) #['s', 'u', 'n', ' ']

#\A与\Z
print(re.findall('\Ahe','hello sun 123')) #['he'],\A==>^
print(re.findall('123\Z','hello sun 123')) #['123'],\Z==>$

# \b与\B
print(re.findall("a\b", 'asbca sdf er dfa')) # 结果['a','a']
print(re.findall("ca\B", 'asbcae sdf er dfa')) # 结果['ca']


# (?P<name>)
print(re.findall("<(?P<tag_name>\w+)>\w+</(?P=tag_name)>","<h1>hello</h1>")) 
    #['h1']
print(re.search("<(?P<tag_name>\w+)>\w+</(?P=tag_name)>","<h1>hello</h1>").group())
    #<h1>hello</h1>
print(re.search("<(?P<tag_name>\w+)>\w+</(?P=tag_name)>","<h1>hello</h1>").groupdict())
    # {'tag_name': 'h1'}
print(re.search("<(?P<tag_name>\w+)>\w+</(?P=tag_name)>","<h1>hello</h1>").groups())
    # ('h1',)
如何使用上述示例
re方法
import re
re.findall("regex", 字符串)
    # 把所有找到的字符串以列表的方式返回

re.search('e','abcdefg').group()#['e']
    # 只到找到第一个匹配然后返回一个包含匹配信息的对象,该对象可以通过调用group()方法得到匹配的字符串,
    # 如果字符串没有匹配,则返回None。
re.match('e','abcde'))
    #None,同search,不过在字符串开始处进行匹配,完全可以用search+^代替match
re.split('[ab]','abcd')
    #['', '', 'cd'],先按'a'分割得到''和'bcd',再对''和'bcd'分别按'b'分割
re.sub('a','A','abcabc')
    # 结果AbcAbc,不指定n,默认替换所有
re.sub('a','A','abcabc', 1)
    # 结果Abcabc
re.subn('a','A','abcabc')
    # 结果('AbcAbc', 2), 结果带有替换的个数

obj=re.compile('\d{2}')
obj.search('abc123eeee').group() # 结果12
obj.findall('abc123eeee') # 结果['12']
re内置方法
#使用|,先匹配的先生效,|左边是匹配小数,而findall最终结果是查看分组,所有即使匹配成功小数也不会存入结果
#而不是小数时,就去匹配(-?\d+),匹配到的自然就是,非小数的数,在此处即整数
#
print(re.findall(r"-?\d+\.\d*|(-?\d+)","1-2*(60+(-40.35/5)-(-4*3))")) #找出所有整数['1', '-2', '60', '', '5', '-4', '3']
findall需要注意的地方
# 现有字符串如下
src = "c++|java|python|shell"
# 用正则表达式将c 和shell换位置

# 先用分组将 内容 分为三个  1.c++  2.|java|python| 3.shell
print(re.findall("(.+?)(\|.+\|)(.+)",src))
print(re.sub("(.+?)(\|.+\|)(.+)",r"\3\2\1",src))

# print(re.search("(.+?)(\|.+)\|(.+)",src).group(3))
# group(3)这里的3和上述的\3y意思一样,代表的是第3部分的分组


# 总结:
# 分组后,如果两个挨近的分组都为贪婪模式,则以第一个贪婪,第二个非贪婪
src = "c++|java|python|php|.net|vue|shell" 
print(re.findall("(.+)(\|.+\|)(.+)",src))
# 结果是:[('c++|java|python|php|.net', '|vue|', 'shell')]
# 3个分组中都为贪婪模式,所以会以第1个贪婪,第2、3个贪婪,
# 如果第1个为非贪婪,则第2个贪婪,第3个非贪婪
实现字符串的反转

 

posted @ 2019-04-06 13:13  yw_sun  阅读(316)  评论(0)    收藏  举报