惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

T
Tor Project blog
AI
AI
S
Securelist
P
Privacy International News Feed
A
Arctic Wolf
T
Tenable Blog
C
Cisco Blogs
P
Proofpoint News Feed
Application and Cybersecurity Blog
Application and Cybersecurity Blog
Google Online Security Blog
Google Online Security Blog
S
Schneier on Security
AWS News Blog
AWS News Blog
L
Lohrmann on Cybersecurity
D
Darknet – Hacking Tools, Hacker News & Cyber Security
N
News and Events Feed by Topic
Know Your Adversary
Know Your Adversary
H
Heimdal Security Blog
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
Cyberwarzone
Cyberwarzone
C
Cybersecurity and Infrastructure Security Agency CISA
S
Security Affairs
P
Palo Alto Networks Blog
K
Kaspersky official blog
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
博客园 - 叶小钗
Recent Commits to openclaw:main
Recent Commits to openclaw:main
博客园 - Franky
SecWiki News
SecWiki News
IT之家
IT之家
G
GRAHAM CLULEY
酷 壳 – CoolShell
酷 壳 – CoolShell
C
CERT Recently Published Vulnerability Notes
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
L
LINUX DO - 最新话题
宝玉的分享
宝玉的分享
月光博客
月光博客
H
Help Net Security
P
Proofpoint News Feed
Cloudbric
Cloudbric
Latest news
Latest news
Spread Privacy
Spread Privacy
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Schneier on Security
Schneier on Security
Help Net Security
Help Net Security
Apple Machine Learning Research
Apple Machine Learning Research
Webroot Blog
Webroot Blog
B
Blog
量子位
J
Java Code Geeks
MyScale Blog
MyScale Blog

晚花行乐

马克卡尼在2026年达沃斯论坛上的讲话(阅读材料) | 晚花行乐 鸡娃如何用力才是恰到好处 | 晚花行乐 读万卷书,行万里路的辩证关系 | 晚花行乐 反对培训机构掐尖招生 | 晚花行乐 小泽和建国会谈最后10分钟全文(阅读材料) | 晚花行乐 来看看 DeepSeek 怎么鸡娃 | 晚花行乐 谈谈基本功 | 晚花行乐 惠普 ProDesk SFF PC 各系列参数对比 | 晚花行乐 解决瘦客户机上安装 Debian 12 启动失败问题 | 晚花行乐 意拾喻言:老外写的文言文 | 晚花行乐 笠翁对韵中的典故(十三元) | 晚花行乐 谈谈中考取消四小门 | 晚花行乐 亚马逊云科技产品免费试用攻略(3) - 对象存储服务 | 晚花行乐 在 Windows 10 LTSC 版本上安装 WSL2 | 晚花行乐 Debian 12 的常用配置项 | 晚花行乐 在 Debian 12 上安装 Nvidia 显卡驱动程序 | 晚花行乐 解决 Debian 12 关机失败问题 | 晚花行乐 解决 VS Code 自动更新版本后卡在连接界面 | 晚花行乐 观看巴黎奥运会有感 | 晚花行乐 在 Windows10 上安装惠普旧打印机驱动程序 | 晚花行乐 欢迎关注公众号:晚花行乐 | 晚花行乐 如何编写拼写检查器 | 晚花行乐 亚马逊云科技产品免费试用攻略(2) - 云服务器 | 晚花行乐 Pandas 中 axis 参数的理解(附实例) | 晚花行乐 我打算命个名,叫什么什么 Manager | 晚花行乐 上海武康路历史建筑一览 | 晚花行乐 Python 实现简单的数学表达式解析并处理 | 晚花行乐 观看马拉松的感悟 | 晚花行乐 Python 保存 Cookies 到文件并再次读取 | 晚花行乐 如何为 Hugo 静态网站添加评论功能 | 晚花行乐 Linux 共享打印服务 CUPS | 晚花行乐 如何为 Hugo 静态网站添加搜索功能 | 晚花行乐 解决 CSV 文件的第一列不能解析 | 晚花行乐 亚马逊云科技产品免费试用攻略(1) - 注册账户 | 晚花行乐 第二幕 Atma 的闲聊 | 晚花行乐 第二幕野蛮人的语音 | 晚花行乐 古入声和普通话平声对照 | 晚花行乐 第一幕的背景音乐 | 晚花行乐 第二幕亚马逊的语音 | 晚花行乐 第一幕的亚马逊的语音 | 晚花行乐 第一幕的野蛮人的语音 | 晚花行乐 MacOS 的彩蛋:Here's to the crazy ones | 晚花行乐 笠翁对韵的基本知识 | 晚花行乐 笠翁对韵中的典故(十二文) | 晚花行乐 笠翁对韵中的典故(十一真) | 晚花行乐 笠翁对韵中的典故(十灰) | 晚花行乐 杭州景点的楹联 | 晚花行乐 adb keycode 大全 | 晚花行乐 Scikit-learn 学习笔记(0)名词术语 | 晚花行乐 Scikit-learn 学习笔记(3)监督学习的例子 | 晚花行乐 SQLite 文档的学习笔记(1)长期支持计划 | 晚花行乐 SQLite 文档的学习笔记(2)测试方法 | 晚花行乐 笠翁对韵中的典故(九佳) | 晚花行乐 Ansible 如何检查一个程序的版本 | 晚花行乐 Ansible 如何检查一个文件夹是否存在 | 晚花行乐 pip 配置文件详解 | 晚花行乐 Ansible 如何检查一个URL是否正常 | 晚花行乐 Ansible 如何修改 iptables 规则 | 晚花行乐 Ansible 指定 playbook 运行的主机 | 晚花行乐 Ansible 如何清空文件夹 | 晚花行乐 Ansible 如何在本机执行命令 | 晚花行乐 笠翁对韵中的典故(八齐) | 晚花行乐 Python 中 Defaultdict 的理解 | 晚花行乐 《伊索寓言》电子书 | 晚花行乐 菲伯尔钢琴伴奏:第二册 | 晚花行乐 Python 的 Keyword-Only Arguments 理解 | 晚花行乐 Python 的 函数参数处理机制 | 晚花行乐 瓦瑞夫在第一幕的闲聊 | 晚花行乐 瓦瑞夫在第一幕的任务提示 | 晚花行乐 第一幕的女巫语音 | 晚花行乐 《Fluent Python》 读书笔记:文本和字节序列 | 晚花行乐 第一幕的罗格语音 | 晚花行乐 第一幕的圣骑士语音 | 晚花行乐 第一幕的男巫语音 | 晚花行乐 第一幕的旁白 | 晚花行乐 第一幕的恶魔 | 晚花行乐 《Fluent Python》 读书笔记:字典和集合 | 晚花行乐 笠翁对韵中的典故(七虞) | 晚花行乐 笠翁对韵中的典故(五微) | 晚花行乐 笠翁对韵中的典故(六鱼) | 晚花行乐 笠翁对韵中的典故(一东) | 晚花行乐 笠翁对韵中的典故(二冬) | 晚花行乐 笠翁对韵中的典故(三江) | 晚花行乐 笠翁对韵中的典故(四支) | 晚花行乐 姜太公钓鱼 | 晚花行乐 武王建立周朝 | 晚花行乐 大禹治水 | 晚花行乐 尧舜让位 | 晚花行乐 黄帝战蚩尤 | 晚花行乐 上下五千年-精简版 | 晚花行乐 成功修复鼠标按键 | 晚花行乐 横向的Word文档怎么加页眉页脚 | 晚花行乐 商标通用化的故事:商标代替商品名 | 晚花行乐 搜索空文件夹的批处理程序 | 晚花行乐 Sn0wbreeze不能运行? | 晚花行乐 天线的驻波比 | 晚花行乐 天线参数:增益Gain | 晚花行乐 天线参数:方向图Radiation pattern | 晚花行乐 本拉登别墅的Google Earth坐标 | 晚花行乐 宜家帕克斯(PAX)衣柜的拼装过程 | 晚花行乐
Python 中 Element Tree 的理解 | 晚花行乐
2021-12-13 · via 晚花行乐

ElementTree XML 模块用来处理 XML 格式文档。包括读取、解析、修改、生成等等。

xml.etree.ElementTree 库里主要使用的是两个类:

ElementTree 是对整个文档树(document tree)的抽象,操作文件、文档时需要。比如读取、写入、查找。

Element 是对文档中单个节点(document node)的抽象,操作单个节点时需要。比如修改属性、查找子节点。

读入 xml 文件

有两种途径,各自得到不同的结果。

返回 xml 树

使用库函数: xml.etree.ElementTree 库内的 parse() 函数。

返回的是一个 ElementTree 树对象

>>> from xml.etree.ElementTree import parse
>>> parse('a.xml')
<xml.etree.ElementTree.ElementTree object at 0x0000020DEAF736D0>
>>>

如果想得到树的根节点,还需要使用 getroot() 函数:

>>> from xml.etree.ElementTree import parse
>>> parse('a.xml').getroot()
<Element 'xml' at 0x0000020DEB462360>
>>>

返回 xml 根节点

使用类方法: xml.etree.ElementTree 库内的 ElementTree 类的实例方法。

可以读入文件名或者文件对象。

例子:从文件名读入:

>>> from xml.etree.ElementTree import ElementTree
>>> ElementTree().parse('a.xml') 
<Element 'xml' at 0x0000020DEB462310>

可以看出,返回的是 Element 对象,这个对象是文档树的根。

例子:从文件对象读入

>>> with open('a.xml','r') as x:
...     ElementTree().parse(x)        
... 
<Element 'xml' at 0x0000020DEB462400>
>>>

从字符串读入

使用库函数: xml.etree.ElementTree 库内的 fromstring() 函数。

>>> import xml.etree.ElementTree as ET
>>> ET.fromstring('<xml></xml>')
<Element 'xml' at 0x000001FD319BDEA0>
>>>

返回根节点。

输出字符串

使用类方法:xml.etree.ElementTree 库内的 dump()tostring()函数。

dump() 向标准输出打印出字符串,函数不返回字符串,仅用于调试。参数是节点对象 Element。

>>> import xml.etree.ElementTree as ET
>>> t=ET.fromstring('<xml><a>1</a></xml>') 
>>> ET.dump(t)
<xml><a>1</a></xml>
>>>

tostring() 用来返回格式化后的 xml字符串,参数列表有:

tostring(element, 
      encoding='us-ascii', 
      method='xml', 
      *
      xml_declaration=None, 
      default_namespace=None, 
      short_empty_elements=True)

参数中,element 是个 Element 节点对象。例子:

>>> import xml.etree.ElementTree as ET
>>> t=ET.fromstring('<xml><a>1</a></xml>') 
>>> s=ET.tostring(t)
>>> s
b'<xml><a>1</a></xml>'
>>>

得到字符串后,就可以编写代码输出到文件。

encoding 参数指定编码,默认是 ascii 的字节编码。如果要输出 UTF-8 编码,可以赋值:“unicode”。例子:

>>> import xml.etree.ElementTree as ET
>>> t=ET.fromstring('<xml><a>1</a></xml>')
>>> s=ET.tostring(t, encoding="unicode")
>>> s
'<xml><a>1</a></xml>'
>>>

xml_declaration 参数控制是否同时输出xml 格式声明。例子:

>>> import xml.etree.ElementTree as ET
>>> t=ET.fromstring('<xml><a>1</a></xml>')
>>> s=ET.tostring(t, encoding="unicode", xml_declaration=True)
>>> s
"<?xml version='1.0' encoding='cp936'?>\n<xml><a>1</a></xml>"
>>> s=ET.tostring(t,  xml_declaration=True)                    
>>> s
b"<?xml version='1.0' encoding='us-ascii'?>\n<xml><a>1</a></xml>"
>>>

或者直接使用 XML 树对象自带的方法:write(),在下节介绍。

写入 xml 文件

函数的参数有:

write(file, 
      encoding='us-ascii', 
      xml_declaration=None, 
      default_namespace=None, 
      method='xml', 
      *, 
      short_empty_elements=True)

file: 和 读入 xml 文件 一样,可以提供文件名或者文件对象。

其他参数和 tostring() 相同。

处理 xml 命名空间

声明命名空间的方法是声明一个 prefix 到 uri 的映射:

xmlns:prefix=uri

这里 prefix 是个临时替代字符串,用于替换 uri。解析 xml 时,所有tag 前面的 prefix 会被自动替换为 uri

如果没有 prefix,而采用下面形式

xmlns:uri

就是默认命名空间,default namespace,所有没有 prefix 的前缀会被自动替换为 uri

如果原始的 xml 声明了命名空间,在解析的时候会自动将扩展tag。

借用 这里的例子 说明命名空间的处理。

下面一个 xml 文件里,同时声明了 html 类型的 table 和家具类型的 table, 然后用命名空间区分。

<?xml version='1.0'?>
<root xmlns:h="html"
       xmlns:f="furniture"
       xmlns="nothing">
  <h:table>
    <h:tr>
    <h:td>Apples</h:td>
    <h:td>Bananas</h:td>
    </h:tr>
  </h:table>
  <f:table>
    <f:name>African Coffee Table</f:name>
    <f:width>80</f:width>
    <f:length>120</f:length>
  </f:table>
</root>

上面文件的例子中,default namespace 声明为 nothing,另外声明了两种 namespace 分别为:

h="html"

f="furniture"

所有 h: 开头的tag都会被替换为 html:,比如h:table会被替换为html:table。同样道理,f:table会被替换为furniture:table。而没有前缀的则会添加 default namespace 的值。

扩展版本的 xml如下:

<?xml version='1.0'?>
<nothing:root xmlns:h="html"
       xmlns:f="furniture"
       xmlns="nothing">
  <html:table>
    <html:tr>
    <html:td>Apples</html:td>
    <html:td>Bananas</html:td>
    </html:tr>
  </html:table>
  <furniture:table>
    <furniture:name>African Coffee Table</furniture:name>
    <furniture:width>80</furniture:width>
    <furniture:length>120</furniture:length>
  </furniture:table>
</nothing:root>

因为 prefix 是中间临时替代品,所以更改 prefix 并不会影响最终扩展出的xml。经过 ElementTree.parse() 后会统一替换 prefix 为统一编号的 ns,从ns0ns1……

>>> import xml.etree.ElementTree as ET
>>> a=ET.ElementTree()
>>> a.parse('a.xml')
>>> ET.dump(a.getroot())
<ns0:root xmlns:ns0="nothing" xmlns:ns1="html" xmlns:ns2="furniture">
<ns1:table>
   <ns1:tr>
   <ns1:td>Apples</ns1:td>
   <ns1:td>Bananas</ns1:td>
   </ns1:tr>
</ns1:table>
<ns2:table>
   <ns2:name>African Coffee Table</ns2:name>
   <ns2:width>80</ns2:width>
   <ns2:length>120</ns2:length>
</ns2:table>

</ns0:root>

当我们查看元素 tag 的名称时,会看到放置在大括号中的 uri:

>>> r=a.getroot()
>>> r.tag
'{nothing}root'

当用 find() 寻找子元素时,需要提供扩展版的 tag 名称, 如下例子:

>>> r.find('{html}table')  
<Element '{html}table' at 0x000002608E572720>

这样并不方便,应该使用 find() 函数的第二个参数:ns,来指定命名空间,形式是从 prefix 到 uri 的映射。比如:

>>> r.find('h:table', {'h':'html'})     
<Element '{html}table' at 0x000002608E572720>

格式化

默认情况下,tostring()write() 的结果将保留原始文件的缩进格式,不会对空行、缩进做统一的格式化。

比如:如果原文件是这样:

<?xml version='1.0'?>
<root xmlns:h="html" xmlns:f="furniture" xmlns="nothing">
<h:table><h:tr><h:td>Apples</h:td><h:td>Bananas</h:td></h:tr></h:table>
<f:table><f:name>African Coffee Table</f:name><f:width>80</f:width><f:length>120</f:length></f:table>
</root>

那么经 ElementTree.parse() 处理后再输出是这样:

>>> import xml.etree.ElementTree as ET
>>> a=ET.ElementTree()
>>> a.parse('a.xml')
<Element '{nothing}root' at 0x0000020A5B822310>
>>> ET.dump(a.getroot())
<ns0:root xmlns:ns0="nothing" xmlns:ns1="html" xmlns:ns2="furniture">
<ns1:table><ns1:tr><ns1:td>Apples</ns1:td><ns1:td>Bananas</ns1:td></ns1:tr></ns1:table>
<ns2:table><ns2:name>African Coffee Table</ns2:name><ns2:width>80</ns2:width><ns2:length>120</ns2:length></ns2:table>
</ns0:root>
>>>

如果想得到如下的两种格式化,就要使用 ElementTree.indent() 函数,这个函数将整棵 xml 文档树重新整理,参数就是文档树的对象 ElementTree。还是上面的例子:

>>> import xml.etree.ElementTree as ET
>>> a=ET.ElementTree()
>>> a.parse('a.xml')
<Element '{nothing}root' at 0x0000020A5B822310>
>>> ET.indent(a)
>>> ET.dump(a.getroot())
<ns0:root xmlns:ns0="nothing" xmlns:ns1="html" xmlns:ns2="furniture">
  <ns1:table>
    <ns1:tr>
      <ns1:td>Apples</ns1:td>
      <ns1:td>Bananas</ns1:td>
    </ns1:tr>
  </ns1:table>
  <ns2:table>
    <ns2:name>African Coffee Table</ns2:name>
    <ns2:width>80</ns2:width>
    <ns2:length>120</ns2:length>
  </ns2:table>
</ns0:root>
>>>

indent() 做了两件事情:

  • 每个节点独占一行
  • 每一级节点增加一级缩进

各位读后有什么想法,请在下方留言吧!如果对本文有疑问或者寻求合作,欢迎 联系邮箱邮箱已到剪贴板

精彩评论