摘要:
Use BeautifulSoupfrom urllib import urlopenfrom bs4 import BeautifulSoup as BStext = urlopen("http://www.python.org/community/jobs/").read()soup = BS(text.decode('gbk', 'ignore'))jobs = set()for header in soup('h2'): links = header('a', 'reference') 阅读全文
posted @ 2012-05-22 22:20
小楼
阅读(257)
评论(0)
推荐(0)
摘要:
Using HTMLPareserUsing HTMLParser simply means subclassing it, and overriding various event-handling methods such as handle_starttag or handle_data.Handle_starttag(tag, attrs): When a start tag is found. Attrs is a sequence of (name, value) pairs.Handle_startendtag(tag, attrs): for empty tags; defau 阅读全文
posted @ 2012-05-22 22:19
小楼
阅读(226)
评论(0)
推荐(0)
摘要:
Screen scraping is a process whereby your program downloads Web pages and extracts information from them. Conceptually, the technique is very simple. You download the data and analyze it, you could, simply use urllib, get the Web page’s HTML source, and then use regular expressions or some such to e 阅读全文
posted @ 2012-05-22 22:18
小楼
阅读(269)
评论(0)
推荐(0)
浙公网安备 33010602011771号