惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

N
News | PayPal Newsroom
IT之家
IT之家
Jina AI
Jina AI
博客园 - 司徒正美
GbyAI
GbyAI
WordPress大学
WordPress大学
B
Blog
大猫的无限游戏
大猫的无限游戏
Y
Y Combinator Blog
阮一峰的网络日志
阮一峰的网络日志
Blog — PlanetScale
Blog — PlanetScale
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Recorded Future
Recorded Future
T
Threat Research - Cisco Blogs
AWS News Blog
AWS News Blog
Latest news
Latest news
宝玉的分享
宝玉的分享
小众软件
小众软件
NISL@THU
NISL@THU
C
CERT Recently Published Vulnerability Notes
The GitHub Blog
The GitHub Blog
P
Privacy & Cybersecurity Law Blog
P
Palo Alto Networks Blog
Spread Privacy
Spread Privacy
Last Week in AI
Last Week in AI
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
P
Proofpoint News Feed
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
量子位
博客园_首页
T
The Exploit Database - CXSecurity.com
The Cloudflare Blog
M
MIT News - Artificial intelligence
H
Help Net Security
Security Archives - TechRepublic
Security Archives - TechRepublic
V2EX - 技术
V2EX - 技术
I
InfoQ
D
Darknet – Hacking Tools, Hacker News & Cyber Security
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
O
OpenAI News
MongoDB | Blog
MongoDB | Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
P
Privacy International News Feed
Microsoft Security Blog
Microsoft Security Blog
C
Cybersecurity and Infrastructure Security Agency CISA
Google DeepMind News
Google DeepMind News
H
Hacker News: Front Page
W
WeLiveSecurity
N
News and Events Feed by Topic

Hi, I Am I

[I Am I 年度简报] — 不知终日梦为鱼 初探 ESP32-CAM QQ 聊天记录 MHT 文件转 HTML [I Am I 年度简报] - 草木本无意,荣枯自有时。 Hexo 中实现 Live Photos 支持 写在当下 NKCTF 2024 1z_F0r3ns1c5 Writeup 春秋杯冬季赛 2023 Writeup [I Am I 年度简报] - 2023 某内网渗透内部赛 Writeup 强网拟态 2023 Writeup Github Actions 自动化部署 Hexo 浅析CobaltStrike流量解密 陇剑杯 2023 Writeup CTF线下赛AWDP总结 ISCC 2023 Writeup ISCC 2023 实战题 Writeup CISCN 2023 Writeup 福建闽盾杯网络空间安全大赛 2023 Writeup 天一永安杯宁波市网络安全大赛 2023 Writeup 贵阳大数据及网络安全精英对抗赛 2023 Writeup 红明谷杯 2023 Writeup Confetti 带来有仪式感的鼓励 记一次 JS 逆向密码加密 [I Am I 年度简报] – 2022 PHP 读取 Excel 文件内容并写入数据库 从0开始的 MoeCTF 开发之路 观安杯 2022 Writeup 利用微信服务号实现早安自动化 Cloudflare批量拉黑IP脚本 蓝帽杯 2022 Writeup 为你的网站添加 Do you like me 小组件 ISCC 2022 Writeup CISCN 2022 Writeup ISCC 2022 实战题 Writeup CTF线下赛AWD攻防总结 [I Am I 年度简报] – 2021 记一次 CNVD 通用型漏洞证书挖掘 Google Adsense 收款流程 使用 Digispark 开发板制作 BadUSB 在 Vue 中使用 Axios 获取钉钉群直播回放的 M3U8 地址 Flask 框架学习记录 关于 Ten·API 防火墙的配置 关于最近 关于 Burp Suite 调教这档事 ISCC 2021 Writeup 记一次 CTF 环境和动态独立靶机部署 各大平台图集解析思路 情话总雷同,恨意千万种 HackThisSite Basic Writeup [I Am I 年度简报] - 2020 Kali 设置中文语言和更换镜像源 使用 Python 下载哔哩哔哩视频 PHP使用 CURL 发送网络请求 PHP蓝奏云直链解析源码 从0开始写一个短视频去水印接口 Windows+Ubuntu 双系统之美化 GRUB Windows+Ubuntu 双系统安装 树莓派安装 Aria2 实现24小时不间断的下载机 抖音无水印解析最新PHP源码 API-Admin Ten·API管理后台 树莓派安装 DLNA 实现流媒体服务器 [I Am I 年度简报] - 2019 Hello Hexo Goindex 将 Google Drive 打造成网盘 自用博客评论邮件通知美化模板 使用 IFTTT 长久保留 Google Voice 号码 Telegram MTProxy 代理一键安装脚本 三分钟学会搭建我的世界基岩版服务器 PHP跳转QQ聊天窗口源码 推荐几款开源HTML5(Web)框架 谷歌新出浏览器 Chromium 可以直接翻墙 Live2D!为你的网站添加看板娘 为你的网站添加Gittalk评论 网站数据离奇丢失...... Lsky Pro(兰空图床)—又一款单纯的图床程序 LOL明天解封ヾ(◍°∇°◍)ノ゙ 唔~本站受到DDOS攻击 [I Am I 年度简报] – 2018 死肥宅也要谈恋爱之早安晚安自动化 1024,Hello,world! 宝塔面板 Bt.cn 专业版破解教程 Uptime >>16 years 震惊!坐在家里竟然可以日入百万 免费获取一年的 .ooo 域名 PHP 调用 新浪API 生成短网址 思杰马克丁成 Adobe 中国授权经销商 畅言商业广告上线-去除畅言评论广告 使用网易云音乐官方接口解析VIP音乐 PHP获取QQ昵称和头像API 坦白说查发送人QQ新方法(已失效) 日常水一波 通过微博图片地址溯源上传者 WordPress 评论夜间自动改为必须审核 十步叫你如何无损修复硬盘锁(mbr病毒) [I Am I 年度简报] – 2017 密码破解与心理学 网页屏蔽各种按键的代码分享 网络安全技术专业术语
从学习通复制文字乱码看前端版权保护
2024-03-01 · via Hi, I Am I

写在前面

起因是在学习通答题的时候突然发现复制出来的内容是乱码,具体示例如下

通过修改HTTP headers 中的哪个键值可以伪造来源网址
嶲嶱修改HTTP headers 中的哪嶮嶭值嶰以嶯造来嶬网址			

分析过程

通过测试,发现仅 章节检测 中存在文字复制乱码,其他答题页面不存在,我们直接通过 devtools 定位到字体文件

image-20240220182057590

@font-face {
  font-family:'font-cxsecret';
  src:url('data:application/font-ttf;charset=utf-8;base64,AAEAAAAMAIAAAwBAQkFTRRuOGNgAAFWcAAA...') format('truetype');
}
.font-cxsecret,
.font-cxsecret p,
.font-cxsecret div,
.font-cxsecret i,
.font-cxsecret em,
.font-cxsecret b,
.font-cxsecret strong,
.font-cxsecret a,
.font-cxsecret font,
.font-cxsecret span,
.font-cxsecret pre,
.font-cxsecret code {
  font-family: 'font-cxsecret' !important;
}

将字体文件存储后,我们发现貌似仅仅是字形名称进行了更改

>>> f"uni{ord('下'):X}"
'uni4E0B'
>>> chr(int('uni5DD5'[3:], 16))
'巕'

根据代码所示, 字的编码应为 uni4E0B,而从学习通下载的字体中 字的编码为 uni5DD5,实际文字应为 。我这里使用的字体查看器是 FontLab,你也可以使用 在线字体编辑器

image-20240308141743656

我们通过获取到的学习通字体信息,检索并下载原版 思源黑体 字体文件

SourceHanSansCN-Normal · Regular · Version 1.000;PS 1;hotconv 1.0.78;makeotf.lib2.5.61930

然后通过 TTFont 库分别获取 学习通字体思源黑体 的字形数据

from fontTools.ttLib import TTFont
# font = TTFont('学习通.ttf')
# font.saveXML('学习通.xml')
font = TTFont('Source Han Sans CN Normal.ttf')
font.saveXML('学习通.xml')

将获取到的字形数据进行对比,我们发现除了 name 不同外,字形数据是一致的

# 学习通字体
<TTGlyph name="uni5DD5" xMin="59" yMin="-72" xMax="942" yMax="762">
  <contour>
    <pt x="515" y="695" on="1"/>
    <pt x="515" y="517" on="1"/>
	...
    <pt x="942" y="762" on="1"/>
    <pt x="942" y="695" on="1"/>
  </contour>
  <instructions/>
</TTGlyph>
# 原版思源黑体
<TTGlyph name="uni4E0B" xMin="59" yMin="-72" xMax="942" yMax="762">
  <contour>
    <pt x="515" y="695" on="1"/>
    <pt x="515" y="517" on="1"/>
	...
    <pt x="942" y="762" on="1"/>
    <pt x="942" y="695" on="1"/>
  </contour>
  <instructions/>
</TTGlyph>

字体解密

那么到此我们就已经理顺了学习通字体的加密思路,简单来说就是以下几个步骤:

  1. 更改字形名称:学习通将原版思源黑体的字形名称进行了更改,例如, 字的编码从 uni4E0B 变为 uni5DD5。这种更改使得直接查看字体文件时,无法直接对应到正确的字符。
  2. 保持字形数据不变:尽管字形名称发生了变化,但是字形数据并没有改变。这意味着,如果我们能够找到字形名称的映射关系,就可以正确地解析出字符。

因此,要解密学习通的字体,我们需要做的就是找到字形名称的映射关系。这可以通过比较学习通字体和原版思源黑体的字形数据来实现。

但是经过一系列测试,发现有些坐井观天了,pt 里面的数据只有 横竖 相同,涉及到 撇捺 的字体,学习通对其进行了一些简单位移

具体如下图所示,左侧为原版字体,右侧为学习通字体

image-20240314184442887

所以我们只能尝试忽略曲线,看看能不能达到预期效果

import json
import hashlib
from fontTools.ttLib import TTFont
from fontTools.pens.basePen import BasePen

class GlyphPen(BasePen):
    def __init__(self, glyphSet):
        BasePen.__init__(self, glyphSet)
        self.points = []

    def _moveTo(self, p):
        self.points.append(("move", p))

    def _lineTo(self, p):
        self.points.append(("line", p))

    def _curveToOne(self, p1, p2, p3):
        # 当遇到曲线时,什么都不做
        pass

    def get_points(self):
        return self.points

# file_path = "Source Han Sans CN Normal.ttf"
file_path = "学习通.ttf"
font = TTFont(file_path)

# 提取字形信息
glyph_set = font.getGlyphSet()
result = {}

# 遍历字体中的所有字形
for name in glyph_set.keys():
    # 使用自定义的 GlyphPen 提取字形轮廓
    pen = GlyphPen(glyph_set)
    glyph = glyph_set[name]
    glyph.draw(pen)

    points_str = json.dumps(pen.get_points())
    md5 = hashlib.md5(points_str.encode()).hexdigest()
    result[name] = md5
    
print(json.dumps(result, indent=4))

此时,比对后发现又出现问题,比对仅仅成功了四分之一,譬如 这类整个字体中撇捺占大部分情况下会匹配失败。

后来,又发现 xMinyMinxMaxyMax 貌似是唯一的,我们直接尝试获取这四个值

for name in glyph_set.keys():
    if name.startswith("uni"):
        glyph = glyph_set[name]
        pen = BoundsPen(glyph_set)
        glyph.draw(pen)

        bounds = pen.bounds
        try:
            md5_input = f"xMin={bounds[0]}, yMin={bounds[1]}, xMax={bounds[2]}, yMax={bounds[3]}"
            md5 = hashlib.md5(md5_input.encode('utf-8')).hexdigest()
            result[name] = md5
        except:
            pass

匹配结果相对比获取 字形轮廓 成功率提升较大,但也仅仅停留在 90% 左右 ,部分字体无法比对出正确字体。

原本以为找到映射关系即可,但不知道是学习通有意为之,还是前面找到的原版字体不对。

至此,整个分析过程以失败结束,虽然结局并不完美,但终究还是收获了一些前端版权保护的思路。

写在后面

或许,两种方法结合又或者通过 selenium 元素截图能达到更好的效果。

另外,附比对时用的代码一份

def compare_json(file1, file2):
    with open(file1, 'r') as f:
        data1 = json.load(f)
    with open(file2, 'r') as f:
        data2 = json.load(f)
    reverse_data1 = {v: k for k, v in data1.items()}
    for k, v in data2.items():
        if v in reverse_data1:
            try:
                print(f"{chr(int(k[3:], 16)), k}: {chr(int(reverse_data1[v][3:], 16)), reverse_data1[v]}")
            except:
                print(f"{k}: {reverse_data1[v]}")

最后,通过检索 Github 项目,发现 TellMeYourWish/chaoxing_solution_of_font_confusion 项目中提供了字体映射表,但目前并不知两个 pkl 文件从何而来。