惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

C
Check Point Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
GbyAI
GbyAI
WordPress大学
WordPress大学
月光博客
月光博客
V
Visual Studio Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Google DeepMind News
Google DeepMind News
H
Help Net Security
MongoDB | Blog
MongoDB | Blog
P
Proofpoint News Feed
博客园 - 司徒正美
S
SegmentFault 最新的问题
Apple Machine Learning Research
Apple Machine Learning Research
Blog — PlanetScale
Blog — PlanetScale
B
Blog
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Microsoft Azure Blog
Microsoft Azure Blog
V
V2EX
L
LangChain Blog
腾讯CDC
T
The Blog of Author Tim Ferriss
量子位

老董笔记

尚硅谷机构在哪?尚硅谷培训怎么样?靠谱吗-互联网IT百科 韩顺平介绍,传智讲师,开办泰牛,入尚硅谷等一系列-互联网IT百科 pandas多重索引标准样式(写入excel有空行)-互联网IT百科 cannot join with no overlapping index names-互联网IT百科 pandas多列变多行(即宽表变长表)melt和stack函数-互联网IT百科 pandas多行转多列(长表变宽表)pivot和unstack-互联网IT百科 Index contains duplicate entries, cannot reshape完美解决-互联网IT百科 single positional indexer is out-of-bounds-互联网IT百科 Can only compare identically-labeled Series objects-互联网IT百科 pandas transform用法详解(多个案例)-互联网IT百科 python四舍五入精确实现-互联网IT百科 pandas的groupby使用apply分组排序-互联网IT百科 index 0 is out of bounds for axis 0 with size 0-互联网IT百科 pandas分组过滤filter函数-互联网IT百科 联想Win10系统如何禁用触摸屏关闭触摸-互联网IT百科 brooks seo教程python教程,brooks seo教程网盘,布鲁seo资源-互联网IT百科 电脑右键文件夹一直转圈电卡死怎么回事-互联网IT百科 施琪嘉的心理成长课(荐)-互联网IT百科 百度SEO公司_SEO推广公司哪家好_SEO外包服务如何选-老董笔记 groupby后agg同1列用多个聚合函数、不同列用不同函数、自定义函数-互联网IT百科 pandas的groupby单列多列分组聚合运算-互联网IT百科 DataFrameGroupBy对象及分组个数、分组大小、组名索引、组数据详解-互联网IT百科 pandas中groupby之Grouper and axis must be same length-互联网IT百科 pandas中groupby的分组原理-互联网IT百科 pandas的groupby的使用详解大全-互联网IT百科 openpyxl单元格自动换行强制换行Alignment(wrapText=True)-互联网IT百科 python教程全套(可就业)-互联网IT百科 联想win10系统CPU显示100%,电脑呼呼响怎么回事-互联网IT百科 如何自制CPU,CPU原理是怎么样的?-互联网IT百科 多款视频制作工具(免费)分享及素材推荐-互联网IT百科
groupby分组计算transform转换返回相同长度序列-互联网IT百科
2022-02-17 · via 老董笔记

  groupby是做分组聚合的,理论上既然分组计算了那么每组会有1个值,这样结果数据的索引长度就减少了,有多少组就代表索引的长度。不过有时候我们分组计算后并不希望减少结果数据的索引长度,比如说有个数据源,里面是不同的班级,求每个班级的学生的最高分,然后新增1列最高分放到原来的表中。

  如果想全面了解分组聚合的场景,可以参考文章pandas之groupby使用详解

  比较常见的思路是先根据班级groupby然后应用聚合函数求出最大值,把这个聚合结果再和原来的搬家表进行关联。不过借助transform()函数可以轻松实现这个过程。

  1、transform()新增1列和源数据索引长度相同

  DataFrameGroupBy对象选取1列(选1列就是SeriesGroupBy对象)来应用transform(func)方法

# -*- coding:UTF-8 -*-
import pandas as pd

df = pd.DataFrame({'class': ['一班', '二班','一班', '二班'],
                   'name':['小明','小王','小张','小李'],
                   'score':[100,9,800,7],
                   })
max_score = df.groupby('class')['score'].transform(max)
df['max_score'] = max_score
print(df)

  class name  score  max_score
0    一班   小明    100        800
1    二班   小王      9          9
2    一班   小张    800        800
3    二班   小李      7          9

  2、DataFrameGroupBy对象应用transform()

  DataFrameGroupBy对象应用transform(func)方法时,其函数func传入的参数是源数据分组后每1组的列,与agg方法特点一样。

# -*- coding:UTF-8 -*-
import pandas as pd

def func(ser_col):
    value = ser_col.max() - ser_col.min()
    return value


df = pd.DataFrame({'class': ['一班', '二班','一班', '二班'],
                   'name':['小明','小王','小张','小李'],
                   'score':[100,9,800,7],
                   })
grouped = df.groupby('class')
df = grouped.transform(func)
print(df)
   score
0    700
1      2
2    700
3      2

  以上代码运行可能会出现警告

  FutureWarning: Dropping invalid columns in DataFrameGroupBy.transform is deprecated. In a future version, a TypeError will be raised. Before calling .transform, select only columns which should be valid for the transforming function.

  这个警告是因为我们传入的函数func内部是做数据运算的,而name列是文本不适合做数值计算,提示你选取有效的列来计算,否则在未来的pandas版本中会报错。这也可以证实上面说的DataFrameGroupBy对象应用transform(func)方法时,其函数func传入的参数是源数据分组后每1组的列。

很赞哦!

python编程网提示:转载请注明来源www.python66.com。
有宝贵意见可添加站长微信(底部),获取技术资料请到公众号(底部)。同行交流请加群 python学习会