惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

The GitHub Blog
The GitHub Blog
T
The Blog of Author Tim Ferriss
S
Schneier on Security
Forbes - Security
Forbes - Security
Cisco Talos Blog
Cisco Talos Blog
月光博客
月光博客
T
Threat Research - Cisco Blogs
I
InfoQ
量子位
NISL@THU
NISL@THU
C
Cisco Blogs
云风的 BLOG
云风的 BLOG
P
Privacy & Cybersecurity Law Blog
The Register - Security
The Register - Security
A
Arctic Wolf
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
AWS News Blog
AWS News Blog
T
Troy Hunt's Blog
M
MIT News - Artificial intelligence
B
Blog
T
Tor Project blog
有赞技术团队
有赞技术团队
Hacker News: Ask HN
Hacker News: Ask HN
Y
Y Combinator Blog
L
LangChain Blog
G
Google Developers Blog
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
酷 壳 – CoolShell
酷 壳 – CoolShell
L
LINUX DO - 热门话题
Schneier on Security
Schneier on Security
Cloudbric
Cloudbric
H
Hacker News: Front Page
C
CERT Recently Published Vulnerability Notes
Google DeepMind News
Google DeepMind News
V
V2EX
T
Tailwind CSS Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
O
OpenAI News
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
博客园 - 叶小钗
宝玉的分享
宝玉的分享
罗磊的独立博客
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Scott Helme
Scott Helme
Recorded Future
Recorded Future
Simon Willison's Weblog
Simon Willison's Weblog
J
Java Code Geeks
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
I
Intezer
美团技术团队

博客园 - xiaoyixy

WeakHashMap相关 Java常见问题[转] MONO,原来你是水中月 Lucene 搜索引擎倒排索引原理 PXE Network Boot and Install Linux over NFS server 让进程在后台运行方法汇总 IBM terminology abstraction Perl Note(2) 亲爱的,我想念你 Perl命令行应用 深入PAM Perl Note(1) Shell技巧 GRUB awk学习 高手好习惯 Shell脚本 用户管理 re notes
用 Spreadsheet::ParseExcel处理中文excel文件
xiaoyixy · 2008-10-29 · via 博客园 - xiaoyixy

需要用的模块:
Spreadsheet::ParseExcel
Unicode::Map
IO-stringy
OLE-Storage_Lite

其中IO-stringy,OLE-Storage_Lite为运行的必要包
Spreadsheet::ParseExcel是解析Excel的必要程序
Unicode::Map为完全支持中文的字符集转换包

perldoc Spreadsheet::ParseExcel 介绍的很详细了
中文 excel 处理需要稍微多做一步
用 Spreadsheet::ParseExcel::FmtUnicode 指定 Unicode_Map 为 CP936

然后直接调用 Parse 方法就行了
返回值是一个关键数组 内容就是 excel 文件的内容

$oBook->{File} 是文件名
$oBook->{SheetCount} 是sheet个数
$oBook->{Worksheet}[0]->{Name} 是第一个 sheet 的名字
$oBook->{Worksheet}[1]->{MaxRow} 是第二个 sheet 的最大行
$oBook->{Worksheet}[2]->{Cells}[1][0]->{Val} 是第三个 sheet 的第二行第一列的值
$oBook->{Worksheet}[2]->{Cells}[1][0]->Value 是第三个 sheet 的第二行第一列的转化后的中文值

use Spreadsheet::ParseExcel;
use Spreadsheet::ParseExcel::FmtUnicode;

my $oExcel = new Spreadsheet::ParseExcel;
my $oCode = "CP936";
my $oFmtJ = Spreadsheet::ParseExcel::FmtUnicode->new(Unicode_Map => $oCode);
my $oBook = $oExcel->Parse($excelFile, $oFmtJ);

然后根据情况 循环sheet 行/row 列/column 处理就行了.

#==================================================

Demo

#!/usr/bin/perl -w

use strict;

use Spreadsheet::ParseExcel;

use Spreadsheet::ParseExcel::FmtUnicode;

use Encode qw /from_to/;

my $oExcel = new Spreadsheet::ParseExcel;

die "You must provide a filename to $0 to be parsed as an Excel file"

  unless @ARGV;

#set for charactor

my $oFmtC = Spreadsheet::ParseExcel::FmtUnicode->new( Unicode_Map => "CP936" );

my $oBook = $oExcel->Parse( $ARGV[0], $oFmtC );

my ( $iR, $iC, $oWkS, $oWkC );

print "FILE :",  $oBook->{File},       "\n";

print "COUNT :", $oBook->{SheetCount}, "\n";

print "AUTHOR:", $oBook->{Author},     "\n"

  if defined $oBook->{Author};

for ( my $iSheet = 0 ; $iSheet < $oBook->{SheetCount} ; $iSheet++ )

{

$oWkS = $oBook->{Worksheet}[$iSheet];

print "--------- SHEET:", $oWkS->{Name}, "\n";

for ( my $iR = $oWkS->{MinRow} ;

 defined $oWkS->{MaxRow} && $iR <= $oWkS->{MaxRow} ;

 $iR++ )

{

for ( my $iC = $oWkS->{MinCol} ;

 defined $oWkS->{MaxCol} && $iC <= $oWkS->{MaxCol} ;

 $iC++ )

{

$oWkC = $oWkS->{Cells}[$iR][$iC];

my $val = $oWkC->Value;

$val =~ s/\s+//ig;

from_to( $val, "CP936", "utf8" );

#print "( $iR , $iC ) =>", $oWkC->Value, "\n" if($oWkC);

print "( $iR , $iC ) =>", $val, "\n" if ($oWkC);

}

}

}