惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Engineering at Meta
Engineering at Meta
量子位
云风的 BLOG
云风的 BLOG
P
Proofpoint News Feed
月光博客
月光博客
A
About on SuperTechFans
博客园 - 聂微东
Spread Privacy
Spread Privacy
B
Blog
NISL@THU
NISL@THU
小众软件
小众软件
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
C
Cyber Attacks, Cyber Crime and Cyber Security
Google DeepMind News
Google DeepMind News
Last Week in AI
Last Week in AI
T
Threatpost
Stack Overflow Blog
Stack Overflow Blog
博客园 - 叶小钗
Cyberwarzone
Cyberwarzone
Scott Helme
Scott Helme
P
Privacy & Cybersecurity Law Blog
C
Cisco Blogs
Cisco Talos Blog
Cisco Talos Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
P
Palo Alto Networks Blog
C
Check Point Blog
O
OpenAI News
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
IT之家
IT之家
T
Threat Research - Cisco Blogs
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Know Your Adversary
Know Your Adversary
T
The Exploit Database - CXSecurity.com
N
News and Events Feed by Topic
P
Privacy International News Feed
B
Blog RSS Feed
Google DeepMind News
Google DeepMind News
H
Heimdal Security Blog
Martin Fowler
Martin Fowler
Schneier on Security
Schneier on Security
Webroot Blog
Webroot Blog
The GitHub Blog
The GitHub Blog
V
Visual Studio Blog
V2EX - 技术
V2EX - 技术
V
Vulnerabilities – Threatpost
博客园 - 司徒正美
C
CERT Recently Published Vulnerability Notes
H
Hacker News: Front Page
PCI Perspectives
PCI Perspectives

博客园 - cleardo

大模型的原理学习(一) 我的测试开发十年之路 命令行安装ipa包 ios设备管理 tomcat远程部署 移动线路的测试方案设计 iOS开发过程中的内存监控 iOS添加图片 iOS交叉编译 iOS日志获取 使用flask开发web应用 Python隔离环境的搭建 Ipa打包并安装到iphone 单元测试的痛点 Graphviz入门 Iphone常用工具 iOS越狱后必装软件 构建iOS交叉编译环境 pycurl库使用详解
大模型的原理学习(二)
cleardo · 2026-02-04 · via 博客园 - cleardo

(一) 介绍

(二) 损失函数

(三) CNN

(四) RNN

(五) Transformer架构

(六) GPT

(七) 幻觉问题

(八) RAG

(九) MCP

(十) Skills

第一篇链接回顾: 大模型的原理学习(一)

在第二篇中,我们开始学习下损失函数和拟合的概念。

拟合与损失函数

如何判断拟合的好不好?

损失函数或者代价函数(损失函数的平均值)

拟合动画

损失函数:

1、 均方误差

均方误差

2、 二分类交叉熵

二分类交叉熵

如何优化?

梯度下降法

目标:使得代价函数尽可能小,那就要求偏导

函数 =
偏导

经过激活函数=
激活函数

① 计算损失函数 ,选用二分类交叉熵 得到

损失函数

② 对损失函数中的 w 和 b 求偏导,偏导的作用代表的就是参数变化最快的方向,设置一个步长(学习率 η)

步长

③ 更新参数:每经过一次计算,得到一个 W 和 b 就带入一下计算损失函数,看看损失函数有没有变小,如果变小了,说明本次 w 和 b 的调整是有效的,就更新 w 和 b

④ 继续重复 ①~③ 步骤,直到损失函数达到最小。

梯度下降

在神经网络中,是如何通过一层层隐藏函数,最终找出最优的输出?

前向传播:根据输入 x 计算出预测值 y

反向传播:计算预测值与实际值的损失函数,进行梯度优化

梯度下降2

链式法则

偏导链式法则

第三篇,我们将开始学习卷积神经网络的原理

卷积神经网络(Convolutional Neural Network, CNN)

前面我们都是以一个输入经过神经网络得到一个输出为例,那么当输入很多的时候,神经网络面临什么问题呢?