python模块之feedparser

转置:https://blog.csdn.net/lilong117194/article/details/77323673

feedparser是python中最常用的RSS程序库,使用它我们可轻松地实现从任何 RSS 或 Atom 订阅源得到标题、链接和文章的条目。

使用:pip install feedparser来安装模块

首先随便找了一段简化的rss:

<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
    <title type="text">博客园_mrbean</title>
    <subtitle type="text">**********************</subtitle>
    <id>uuid:32303acf-fb5f-4538-a6ba-7a1ac4fd7a58;id=8434</id>
    <updated>2014-05-14T15:13:36Z</updated>
    <author>
        <name>mrbean</name>
        <uri>http://www.cnblogs.com/MrLJC/</uri>
    </author>
    <generator>feed.cnblogs.com</generator>
    <entry>
        <id>http://www.cnblogs.com/MrLJC/p/3715783.html</id>
        <title type="text">用python读写excel(xlrd、xlwt) - mrbean</title>
        <summary type="text">最近需要从多个excel表里面用各种方式整...</summary>
        <published>2014-05-08T16:25:00Z</published>
        <updated>2014-05-08T16:25:00Z</updated>
        <author>
            <name>mrbean</name>
            <uri>http://www.cnblogs.com/MrLJC/</uri>
        </author>
        <link rel="alternate" href="http://www.cnblogs.com/MrLJC/p/3715783.html" />
        <link rel="alternate" type="text/html" href="http://www.cnblogs.com/MrLJC/p/3715783.html" />
        <content type="html">最近需要从多个excel表里面用各种方式整理一些数据,虽然说原来用过java做这类事情,但是由于最近在学python,所以当然就决定用python尝试一下了。发现python果然简洁很多。这里简单记录一下。(由于是用到什么学什么,所以不算太深入,高手勿喷,欢迎指导)一、读excel表读excel要用...&lt;img src="http://counter.cnblogs.com/blog/rss/3715783" width="1" height="1" alt=""/&gt;&lt;br/&gt;&lt;p&gt;本文链接:&lt;a href="http://www.cnblogs.com/MrLJC/p/3715783.html" target="_blank"&gt;用python读写excel(xlrd、xlwt)&lt;/a&gt;,转载请注明。&lt;/p&gt;</content>
    </entry>
</feed>

把他复制到一个.txt文件中,保存为.xml

 

import feedparser

print feedparser.parse('')
d=feedparser.parse("tt.xml")
print d['feed']['title']  
print d.feed.title        # 通过属性访问
print d.entries[0].id
print d.entries[0].content
-结果:

{'feed': {}, 'encoding': u'utf-8', 'bozo': 1, 'version': u'', 'namespaces': {}, 'entries': [], 'bozo_exception': SAXParseException('no element found',)}
博客园_mrbean
博客园_mrbean
http://www.cnblogs.com/MrLJC/p/3715783.html
[{'base': u'', 'type': u'text/html', 'value': u'\u6700\u8fd1\u9700\u8981\u4ece\u591a\u4e2aexcel\u8868\u91cc\u9762\u7528\u5404\u79cd\u65b9\u5f0f\u6574\u7406\u4e00\u4e9b\u6570\u636e\uff0c\u867d\u7136\u8bf4\u539f\u6765\u7528\u8fc7java\u505a\u8fd9\u7c7b\u4e8b\u60c5\uff0c\u4f46\u662f\u7531\u4e8e\u6700\u8fd1\u5728\u5b66python\uff0c\u6240\u4ee5\u5f53\u7136\u5c31\u51b3\u5b9a\u7528python\u5c1d\u8bd5\u4e00\u4e0b\u4e86\u3002\u53d1\u73b0python\u679c\u7136\u7b80\u6d01\u5f88\u591a\u3002\u8fd9\u91cc\u7b80\u5355\u8bb0\u5f55\u4e00\u4e0b\u3002\uff08\u7531\u4e8e\u662f\u7528\u5230\u4ec0\u4e48\u5b66\u4ec0\u4e48\uff0c\u6240\u4ee5\u4e0d\u7b97\u592a\u6df1\u5165\uff0c\u9ad8\u624b\u52ff\u55b7\uff0c\u6b22\u8fce\u6307\u5bfc\uff09\u4e00\u3001\u8bfbexcel\u8868\u8bfbexcel\u8981\u7528...<img alt="" height="1" src="http://counter.cnblogs.com/blog/rss/3715783" width="1" /><br /><p>\u672c\u6587\u94fe\u63a5\uff1a<a href="http://www.cnblogs.com/MrLJC/p/3715783.html" target="_blank">\u7528python\u8bfb\u5199excel\uff08xlrd\u3001xlwt\uff09</a>\uff0c\u8f6c\u8f7d\u8bf7\u6ce8\u660e\u3002</p>', 'language': None}]
**********************
feedparser 最为核心的函数自然是 parse() 解析 URL 地址的函数,返回的形式如:

{'feed': {}, 'encoding': u'utf-8', 'bozo': 1, 'version': u'', 'namespaces': {}, 'entries': [], 'bozo_exception': SAXParseException('no element found',)}

 
附:spyder ipython 中文乱码的问题

import sys
reload(sys)
sys.setdefaultencoding('utf8')

posted @ 2019-07-25 13:26  Awakenedy  阅读(422)  评论(0)    收藏  举报