1.1如何解压序列赋值给多个变量:如何将列表中的值,赋值给若干变量?

需要保证若干变量的数量和序列中的元素数量一样,如果不匹配会产生异常:

data=['jack',50,91.1,(2020,12,21)]
name,age,price,date=data
print(name,age,price,date)
jack 50 91.1 (2020, 12, 21)

如果只想取其中一部分的值,其他值不需要,可以用占位符_来代替无关值:

s='hello'
_,a,_,_,b=s
print(a,b)
e o

 

1.2如果一个序列元素非常多,如何赋值给其他变量?手动一个个赋值不现实。

可以使用*号来取出对应的若干个值(同理*args在函数参数中即表示若干个参数):

record=('Mike','mike@gmail.com','13753500861','997-332-2322','18322133324')
name,email,*phonenumbers=record #表示序列之后的所有值
print(phonenumbers)
['13753500861', '997-332-2322', '18322133324']
*a,last=[1,2,3,4,5,6,7,8] #a表示最后一个值之前的所有值
print(a)
[1, 2, 3, 4, 5, 6, 7]
t=[1,2,3,4,5,6,7,8,9]
first,*middle,last=t
sum(middle) #计算排除首位元素后的和
35
records = [('foo', 1, 2), ('bar', 'hello'), ('foo', 3, 5)]


def do_foo(x, y):
    print('foo', x, y)


def do_bar(s):
    print('bar', s)


for tag, *args in records:  # tag表示每一个元组的第一个元素,args表示剩下的元素
    if tag == 'foo':
        do_foo(*args)
    elif tag == 'bar':
        do_bar(*args)

*号操作符在操作字符时也很有用,比如分割字符串:

line = 'nobody:*:-2:-2:Unprivileged User:/var/empty:/usr/bin/false'
uname, *field, homedir, sh = line.split(':')
print(uname)
print(homedir)
print(sh)
print(field)

 

1.3如何保留最后的N个元素?

使用collection.deque,deque构造函数会创建一个固定大小的队列。当新元素加入时,老元素会被踢出。

from collections import deque


def search(lines, pattern, history=5):  # lines为待搜索文件每行,pattern为关键字,history为设置保留历史记录行
    previous_lines = deque(maxlen=history)  # 当前存储设为5行
    for line in lines:  # 对于文件每一行
        if pattern in line:  # 如果关键字在此行
            # print(previous_lines)
            # print('*********************')
            yield line, previous_lines  # 返回此行内容,返回当前记录
        previous_lines.append(line)  # 不论关键字,将一行内容追加到previous_lines中


if __name__ == '__main__':
    with open(r'1.txt') as f:
        for line, prevlines in search(f, 'python', 5):
            for pline in prevlines:
                print(pline, end='')
                print(line, end='')
                print('-' * 20)

 

1.4查找最大或最小的N个元素?

heapq模块的两个函数nlargest()和nsmallest()可以完美解决这个问题:

import heapq
nums=[1,8,2,33,23,7,-5,232,42,37,3]
print(heapq.nlargest(3,nums))
[232, 42, 37]
print(heapq.nsmallest(3,nums))
[-5, 1, 2]

 

1.5如何实现优先级队列?并且在这个队列上每次pop操作总返回优先级最高的那个元素

import heapq


class PriorityQueue:
    def __init__(self):
        self._queue = []
        self._index = 0

    def push(self, item, priority):
        heapq.heappush(self._queue, (-priority, self._index, item))  # 优先级为负的目的:使得元素按照优先级从高到低
        self._index += 1

    def pop(self):
        return heapq.heappop(self._queue)[-1]  # 从底部弹出元素


class Item:
    def __init__(self, name):
        self.name = name

    def __repr__(self):
        return 'Item({!r})'.format(self.name)


if __name__ == '__main__':
    q = PriorityQueue()
    q.push(Item('foo'), 1)
    q.push(Item('bar'), 5)
    q.push(Item('spam'), 4)
    q.push(Item('grok'), 1)
    for _ in range(4):
        print(q.pop())

 

1.6怎样实现字典键映射多个值?一个键对应多个值的字典?

如果想要一个键映射多个值,可以将多个值放到另外的容器中,比如一个列表或集合。如:

d={'a':[1,2,3],'b':[4,5]}

可以使用collection模块的defaultdict来构建字典:

from collections import defaultdict

d = defaultdict(list)
d['a'].append(1)
d['a'].append(2)
d['b'].append(4)

print(d)

d = defaultdict(set)
d['a'].add(1)
d['a'].add(2)
d['b'].add(4)

print(d)

 

1.7字典排序,创建一个字典,并在迭代或序列化这个字典时能够控制元素的顺序。

字典默认是无序的。

可以使用collections模块中的OrderDict类,保持元素被插入的顺序:

from collections import OrderedDict

d = OrderedDict()
d['foo'] = 1
d['bar'] = 2
d['spam'] = 3
d['grok'] = 4
for key in d:
    print(key, d[key])

 

1.8怎么样在数据字典中执行一些计算操作(如求最大值,最小值,排序等)

为了对字典执行计算,通常需要使用zip()函数将键和值翻转过来。比如:

prices = {'acme': 45.23, 'APPE': 333.11, 'IBM': 322.1, 'HQ': 66}
min_price = min(zip(prices.values(), prices.keys()))
print(min_price)

 

1.9查找两个字典的相同点:如相同的键或者相同的值等。

a = {'x': 1, 'y': 2, 'z': 3}
b = {'w': 10, 'x': 11, 'y': 2}
print(a.keys() & b.keys())  # 打印出相同的键
print(a.items() & b.items())  # 打印出相同的键值对

 

1.10删除序列相同的元素并保持顺序

def dedupe(items):
    seen = set()
    for item in items:
        if item not in seen:
            yield item
            seen.add(item)


a = [1, 5, 2, 1, 9, 1, 5, 10]
print(list(dedupe(a)))

 

1.11对大量的硬编码下标,建议使用slice命名切片,使得代码更具可读性。

items = [0, 1, 2, 3, 4, 5, 6]
a = slice(2, 4)
print(items[a])

 

1.12找出序列中出现次数最多的元素?

使用collection的Counter函数

words = ['look', 'age', 'hello', 'loob', 'look', 'hello', 'look']
from collections import Counter

word_counts = Counter(words) #统计出现次数

top_one = word_counts.most_common(1) #显示出现次数最多的元素
print(top_one)

 

1.13通过某个关键字排序一个字典列表

通过operator模块的itemgetter函数。

from operator import itemgetter

rows = [{'fname': 'jack', 'lname': 'time', 'uid': 101}, {'fname': 'mike', 'hhah': 'sper', 'uid': 103}, \
        {'fname': 'jones', 'a': 'b', 'uid': 99}, {'fname': 'beazley', 'aaa': 'bbb', 'uid': 200}]
rows_by_fname = sorted(rows, key=itemgetter('fname'))  # 按照fname排序
print(rows_by_fname)

rows_by_uid = sorted(rows, key=itemgetter('uid'))  # 按照uid排序
print(rows_by_uid)

 

1.14如何排序不支持原生比较的对象?

内置的sorted()函数有一个关键字参数key,可以给它传入一个可调用对象,这个对象对每个传入的对象返回一个值用于排序:

class User:
    def __init__(self, user_id):
        self.user_id = user_id

    def __repr__(self):
        return 'User({})'.format(self.user_id)


def sort_notcompare():
    users = [User(23), User(3), User(99)]
    print(users)
    print(sorted(users, key=lambda u: u.user_id))

if __name__ == '__main__':
    sort_notcompare()

 

1.15通过某个字段将记录分组

使用itertools.groupby()函数

 

1.16过滤序列元素

可以通过列表推导式:

mylist = [1, -1, 2, 54, 3, 22, -55, 6]
a = []
[a.append(n) for n in mylist if n > 0]
print(a)

如果堆内存比较敏感,可以使用生成器表达式

 

1.17从字典提取子集

构造一个字典,他是另外字典的子集:

prices = {'ACME': 45.23, 'APPLE': 554.22, 'IBM': 232, 'FB': 111}
p1 = {key: value for key, value in prices.items() if value > 100}
print(p1)

 

1.18映射名称到序列元素

使用collections.nametuple()函数:

from collections import namedtuple

Subscriber = namedtuple('Subscriber', ['addr', 'joined'])  # 生成固定代号元组,可以给元组每个元素命名

sub = Subscriber('test@gmail.com', '2012-1-1')

print(sub)
print(sub.addr)
print(sub.joined)

 

1.19转换并同时计算数据

nums = [1, 2, 3, 4, 5]
s = sum(x * x for x in nums)
print(s)  # 1+4+9+16+25

 

1.20合并多个字典或者映射

collections模块的ChainMap类:

from collections import ChainMap

a = {'x': 1, 'z': 3}
b = {'y': 2, 'z': 4}
c = ChainMap(a, b)
print(c['x'])
print(c['y'])
print(c['z'])  # 先从a中找,找不到再从b中找