pywin32 pywin32 docx文档转html页面 word doc docx 提取文字 图片 html 结构

https://blog.csdn.net/X21214054/article/details/78873338

# python docx文档转html页面 - 程序猿tx - 博客园 https://www.cnblogs.com/taixiang/p/9978456.html
# Usage — PyDocX dev documentation https://pydocx.readthedocs.io/en/latest/usage.html



pywin32 · PyPI https://pypi.org/project/pywin32/


from win32com import client as wc

f = files
# https://docs.microsoft.com/zh-cn/office/dev/add-ins/reference/requirement-sets/office-add-in-requirement-sets?view=office-js

for f in files:
i=10
try:
word = wc.Dispatch('Word.Application')
doc = word.Documents.Open(f)
nf = f.replace('.doc', '.html')
doc.SaveAs(nf, i, False, '', True, '', False, False, False, False) # 创建或覆盖
doc.Close()
word.Quit()
del word, doc # 否则只有一个文件创建
except Exception as e:
print(i, e)

















posted @ 2018-10-06 10:28  papering  阅读(705)  评论(0)    收藏  举报