惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

C
Cybersecurity and Infrastructure Security Agency CISA
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Latest news
Latest news
L
LINUX DO - 热门话题
Cisco Talos Blog
Cisco Talos Blog
S
Securelist
T
Threatpost
AWS News Blog
AWS News Blog
P
Privacy & Cybersecurity Law Blog
C
CERT Recently Published Vulnerability Notes
B
Blog RSS Feed
T
Threat Research - Cisco Blogs
P
Proofpoint News Feed
T
Tor Project blog
P
Palo Alto Networks Blog
博客园 - 三生石上(FineUI控件)
人人都是产品经理
人人都是产品经理
M
MIT News - Artificial intelligence
云风的 BLOG
云风的 BLOG
H
Help Net Security
小众软件
小众软件
C
Cisco Blogs
有赞技术团队
有赞技术团队
Cyberwarzone
Cyberwarzone
雷峰网
雷峰网
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Apple Machine Learning Research
Apple Machine Learning Research
S
Schneier on Security
The GitHub Blog
The GitHub Blog
Y
Y Combinator Blog
The Register - Security
The Register - Security
Project Zero
Project Zero
Hugging Face - Blog
Hugging Face - Blog
The Cloudflare Blog
V
Vulnerabilities – Threatpost
Security Latest
Security Latest
爱范儿
爱范儿
A
About on SuperTechFans
T
The Exploit Database - CXSecurity.com
P
Privacy International News Feed
A
Arctic Wolf
大猫的无限游戏
大猫的无限游戏
V
V2EX
Stack Overflow Blog
Stack Overflow Blog
K
Kaspersky official blog
Scott Helme
Scott Helme
Spread Privacy
Spread Privacy
The Hacker News
The Hacker News
H
Hackread – Cybersecurity News, Data Breaches, AI and More

夜行人

回家路上 第一期的直播演示项目 震动检测器 正能量 在线参观CodeLab Neverland 发布 CodeLab Adapter 3.3.1 DynamicTable 之 纸糊方向盘 CodeLab DynamicTable: 一个可实施的技术方案 CodeLab Insight 发布 Alpha 版 情人节 Home Assistant 周报 && IoT 周报 (02) Joplin: 关注隐私的 Evernote 开源替代软件 浏览器的未来与 Web 传感器 Home Assistant 周报 && IoT 周报 (01) 百宝箱(01) 论自由 介绍 WebThings Home Assistant 周报 && iot 周报 (00) 百宝箱(00) 毛姆读书心得 传世之作 周末徒步 CodeLab Adapter ❤️ Jupyter/Python 航班 躲雨 夏令营途中 [译]思想--作为一种技术 The future of coding 美国之行 三门问题的程序模拟 从Python转向Pharo https://blog.just4fun.site/post/iot/iot-open-source-projects/ Python异步编程笔记 https://blog.just4fun.site/post/iot/iot-open-source-hardware-community/ 万物积木化开发者社区 CodeLab ❤️ Blender Scratch3技术分析之云变量 API(第7篇) [译]对管道(Pipes)的偏爱 [译]提出正确的问题比得到正确答案更重要 蓝牙设备与Scratch3.0 创建你的第一个Scratch3.0 Extension Scratch3技术分析之项目内部数据(第6篇) Scratch3技术分析之社区 API(第5篇) Scratch3技术分析之User API(第4篇) Scratch3技术分析之项目主页API(第3篇) Scratch3技术分析之静态资源API(第2篇) Scratch3.0、micro:bit与Windows7 https://blog.just4fun.site/post/iot/zerynth-vs-micropython/ 核聚变、方所与半宅空间 可视化编程为何是个糟糕的主意 codelab.club周末聚会 关于codelab.club '下一件大事'是一个房间 Hungry Robot - Eat everything 编程作为一种思考方式 今日简史 史蒂夫·乔布斯传 罗素自选文集 https://blog.just4fun.site/post/edx/tianjin-scratch-ai/ https://blog.just4fun.site/post/edx/richie-cms-openedx/ 徒步武功山 WebUSB与micro:bit 积木化编程与3D场景 夜宿武功山顶 scratch3-adapter接入优必选Alpha系列机器人 https://blog.just4fun.site/post/edx/video-migration-note/ scratch3-adapter重构笔记 https://blog.just4fun.site/post/edx/edx-community-members/ 两种硬件编程风格的比较 使用micro:bit自制PPT翻页笔 柏拉图对话集 scratch3.0 + micro:bit 七月电影放映计划 非营利组织的管理 Screenly--用树莓派让任何屏幕变为可编程的数字标牌 以最佳实践开始你的Django项目 micro:bit与事件驱动 为Scratch3.0设计的插件系统(上篇) OCR应用一例 近两年读过的一些好书 blockly开发之使用python驱动浏览器中的turtle(2) 牛顿新传 文学理论入门 逻辑的引擎 人生的意义 blockly开发之生成并运行js代码(1) blockly开发之hello world(0) micro:bit使用笔记 神器之Termux https://blog.just4fun.site/post/iot/micropython-notes/ Cozmo what is this Scratch的前世今生 下段旅程 我行在远方 爆裂 途中杂记 https://blog.just4fun.site/post/edx/open-edx-startup/ cozmo系列之入门 - 有性格且可编程的机器人 PaperWeekly开发笔记 创业二三事
python算法学习之推荐系统
2014-01-07 · via 夜行人

之前一直对算法不太感冒,感觉既乏味又务虚,除了用来考试/面试,实在找不出其他用途。毕竟平时实际项目中也不常遇到需要深入理解算法的地方。加上教科书的影响,对算法一直敬而远之。
直到开始阅读《集体智慧编程》。
才发现原来这也挺好玩的呀。
遂决定好好学习。

长期以来,受惠于推荐系统,日久生情,对此产生了兴趣。比如豆瓣能根据你的浏览记录(评分记录)推荐你可能喜欢的小组,可能喜欢的文章,可能感兴趣的活动。而且往往还真能猜准~
无觅阅读的文章推荐也很和我胃口,不必花太多时间就能轻松找到喜欢的文章。
这类系统能分析出你的品味,它居然知道我的品味!!想想都令人兴奋,系统居然像你的知音一样知道你的品味!!

如果不觉得一件东西有趣,实在很难硬着头皮学下去。既然发现它很好玩,想无视它,也难了。
那就开始我们的算法旅程吧。
从推荐系统开始~

这组文章偏向于总结吧,不是作为入门指南,如果你也对这类算法感兴趣,推荐去阅读《集体智慧编程》,而不是看博客。 这组文章使用的算法皆来自《集体智慧编程》,我只是做些摘录,为了方便日后使用.你也可以使用 这些代码,至于使用过程中你要受到哪些限制,请参考原书的申明部分。

###实例学习 在这个实例里,我们想知道用户的兴趣偏好(口味)。
情景是这样的:几位用户看过几部电影,他们对这些电影进行了评分,我们拿到了这组数据,r如下:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
	critics={'Lisa Rose': {'Lady in the Water': 2.5, 'Snakes on a Plane': 3.5,
	 'Just My Luck': 3.0, 'Superman Returns': 3.5, 'You, Me and Dupree': 2.5, 
	 'The Night Listener': 3.0},
	'Gene Seymour': {'Lady in the Water': 3.0, 'Snakes on a Plane': 3.5, 
	 'Just My Luck': 1.5, 'Superman Returns': 5.0, 'The Night Listener': 3.0, 
	 'You, Me and Dupree': 3.5}, 
	'Michael Phillips': {'Lady in the Water': 2.5, 'Snakes on a Plane': 3.0,
	 'Superman Returns': 3.5, 'The Night Listener': 4.0},
	'Claudia Puig': {'Snakes on a Plane': 3.5, 'Just My Luck': 3.0,
	 'The Night Listener': 4.5, 'Superman Returns': 4.0, 
	 'You, Me and Dupree': 2.5},
	'Mick LaSalle': {'Lady in the Water': 3.0, 'Snakes on a Plane': 4.0, 
	 'Just My Luck': 2.0, 'Superman Returns': 3.0, 'The Night Listener': 3.0,
	 'You, Me and Dupree': 2.0}, 
	'Jack Matthews': {'Lady in the Water': 3.0, 'Snakes on a Plane': 4.0,
	 'The Night Listener': 3.0, 'Superman Returns': 5.0, 'You, Me and Dupree': 3.5},
	'Toby': {'Snakes on a Plane':4.5,'You, Me and Dupree':1.0,'Superman Returns':4.0}}

这些数据是python的字典格式,你如果熟悉js,会发现它几乎就是json格式.
大多网站api接口返回的格式都是json,这样你就知道用python处理这些数据是多么容易.

现在我想挖掘一下这些数据,看看哪两个人的品味比较接近。
摆在我面前的问题是如何度量两个人的品味相似程度呢.毕竟口味这东西不是空间距离可以度量。
距离真是个很好的隐喻。两个品味不同的人就像来自两个星球
如果我们能找到一个数值来度量两个人品味的距离,那么这个问题就变成了数值计算问题!!
还真有这样的东西,我们可以用欧吉里德距离皮尔逊相关度来度量两者相似度.
这里有一个很有趣的概念,叫偏好空间。我们使用以上数据作图:
偏好空间 两个人在偏好空间中距离越近,表示品味越近.如果有多项评分,那么这张图就对应多维。依然适用

我们直接给出计算相似度的欧吉里德方法吧:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
	from math import sqrt

	# Returns a distance-based similarity score for person1 and person2
	def sim_distance(prefs,person1,person2):
	  # Get the list of shared_items
	  si={}
	  for item in prefs[person1]: 
	    if item in prefs[person2]: si[item]=1

	  # if they have no ratings in common, return 0
	  if len(si)==0: return 0

	  # Add up the squares of all the differences
	  sum_of_squares=sum([pow(prefs[person1][item]-prefs[person2][item],2) 
	                      for item in prefs[person1] if item in prefs[person2]])

	  return 1/(1+sum_of_squares)

这里的核心公式,就是数学中简单的两点间距离计算公式,代码只是对此的一个实现,理解时建议在旁边写个数学距离计算公式,那样比代码更能清晰表达算法的本质。

顺便把皮尔逊计算方法也写上,稍后解释:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
	# Returns the Pearson correlation coefficient for p1 and p2
	def sim_pearson(prefs,p1,p2):
	  # Get the list of mutually rated items
	  si={}
	  for item in prefs[p1]: 
	    if item in prefs[p2]: si[item]=1

	  # if they are no ratings in common, return 0
	  if len(si)==0: return 0

	  # Sum calculations
	  n=len(si)
	  
	  # Sums of all the preferences
	  sum1=sum([prefs[p1][it] for it in si])
	  sum2=sum([prefs[p2][it] for it in si])
	  
	  # Sums of the squares
	  sum1Sq=sum([pow(prefs[p1][it],2) for it in si])
	  sum2Sq=sum([pow(prefs[p2][it],2) for it in si])	
	  
	  # Sum of the products
	  pSum=sum([prefs[p1][it]*prefs[p2][it] for it in si])
	  
	  # Calculate r (Pearson score)
	  num=pSum-(sum1*sum2/n)
	  den=sqrt((sum1Sq-pow(sum1,2)/n)*(sum2Sq-pow(sum2,2)/n))
	  if den==0: return 0

	  r=num/den

	  return r

waiting

文章作者 种瓜

上次更新 2014-01-07