Python实现简单HTML表格解析的方法


Posted in Python onJune 15, 2015

本文实例讲述了Python实现简单HTML表格解析的方法。分享给大家供大家参考。具体分析如下:

这里依赖libxml2dom,确保首先安装!导入到你的脚步并调用parse_tables() 函数。

1. source = a string containing the source code you can pass in just the table or the entire page code

2. headers = a list of ints OR a list of strings
If the headers are ints this is for tables with no header, just list the 0 based index of the rows in which you want to extract data.
If the headers are strings this is for tables with header columns (with the tags) it will pull the information from the specified columns

3. The 0 based index of the table in the source code. If there are multiple tables and the table you want to parse is the third table in the code then pass in the number 2 here

It will return a list of lists. each inner list will contain the parsed information.

具体代码如下:

#The goal of table parser is to get specific information from specific
#columns in a table.
#Input: source code from a typical website
#Arguments: a list of headers the user wants to return
#Output: A list of lists of the data in each row
import libxml2dom
def parse_tables(source, headers, table_index):
  """parse_tables(string source, list headers, table_index)
    headers may be a list of strings if the table has headers defined or
    headers may be a list of ints if no headers defined this will get data
    from the rows index.
    This method returns a list of lists
    """
  #Determine if the headers list is strings or ints and make sure they
  #are all the same type
  j = 0
  print 'Printing headers: ',headers
  #route to the correct function
  #if the header type is int
  if type(headers[0]) == type(1):
    #run no_header function
    return no_header(source, headers, table_index)
  #if the header type is string
  elif type(headers[0]) == type('a'):
    #run the header_given function
    return header_given(source, headers, table_index)
  else:
    #return none if the headers aren't correct
    return None
#This function takes in the source code of the whole page a string list of
#headers and the index number of the table on the page. It returns a list of
#lists with the scraped information
def header_given(source, headers, table_index):
  #initiate a list to hole the return list
  return_list = []
  #initiate a list to hold the index numbers of the data in the rows
  header_index = []
  #get a document object out of the source code
  doc = libxml2dom.parseString(source,html=1)
  #get the tables from the document
  tables = doc.getElementsByTagName('table')
  try:
    #try to get focue on the desired table
    main_table = tables[table_index]
  except:
    #if the table doesn't exits then return an error
    return ['The table index was not found']
  #get a list of headers in the table
  table_headers = main_table.getElementsByTagName('th')
  #need a sentry value for the header loop
  loop_sentry = 0
  #loop through each header looking for matches
  for header in table_headers:
    #if the header is in the desired headers list 
    if header.textContent in headers:
      #add it to the header_index
      header_index.append(loop_sentry)
    #add one to the loop_sentry
    loop_sentry+=1
  #get the rows from the table
  rows = main_table.getElementsByTagName('tr')
  #sentry value detecting if the first row is being viewed
  row_sentry = 0
  #loop through the rows in the table, skipping the first row
  for row in rows:
    #if row_sentry is 0 this is our first row
    if row_sentry == 0:
      #make the row_sentry not 0
      row_sentry = 1337
      continue
    #get all cells from the current row
    cells = row.getElementsByTagName('td')
    #initiate a list to append into the return_list
    cell_list = []
    #iterate through all of the header index's
    for i in header_index:
      #append the cells text content to the cell_list
      cell_list.append(cells[i].textContent)
    #append the cell_list to the return_list
    return_list.append(cell_list)
  #return the return_list
  return return_list
#This function takes in the source code of the whole page an int list of
#headers indicating the index number of the needed item and the index number
#of the table on the page. It returns a list of lists with the scraped info
def no_header(source, headers, table_index):
  #initiate a list to hold the return list
  return_list = []
  #get a document object out of the source code
  doc = libxml2dom.parseString(source, html=1)
  #get the tables from document
  tables = doc.getElementsByTagName('table')
  try:
    #Try to get focus on the desired table
    main_table = tables[table_index]
  except:
    #if the table doesn't exits then return an error
    return ['The table index was not found']
  #get all of the rows out of the main_table
  rows = main_table.getElementsByTagName('tr')
  #loop through each row
  for row in rows:
    #get all cells from the current row
    cells = row.getElementsByTagName('td')
    #initiate a list to append into the return_list
    cell_list = []
    #loop through the list of desired headers
    for i in headers:
      try:
        #try to add text from the cell into the cell_list
        cell_list.append(cells[i].textContent)
      except:
        #if there is an error usually an index error just continue
        continue
    #append the data scraped into the return_list    
    return_list.append(cell_list)
  #return the return list
  return return_list

希望本文所述对大家的Python程序设计有所帮助。

Python 相关文章推荐
python遍历序列enumerate函数浅析
Oct 17 Python
基于python中pygame模块的Linux下安装过程(详解)
Nov 09 Python
解决Pycharm界面的子窗口不见了的问题
Jan 17 Python
用Python解决x的n次方问题
Feb 08 Python
Python企业编码生成系统总体系统设计概述
Jul 26 Python
python处理自动化任务之同时批量修改word里面的内容的方法
Aug 23 Python
在Python中获取操作系统的进程信息
Aug 27 Python
Flask框架请求钩子与request请求对象用法实例分析
Nov 07 Python
Python爬虫抓取指定网页图片代码实例
Jul 24 Python
python Paramiko使用示例
Sep 21 Python
python使用torch随机初始化参数
Mar 22 Python
如何利用python创作字符画
Jun 25 Python
Python判断Abundant Number的方法
Jun 15 #Python
Python计算一个文件里字数的方法
Jun 15 #Python
Python素数检测实例分析
Jun 15 #Python
Python计算三维矢量幅度的方法
Jun 15 #Python
Python栈类实例分析
Jun 15 #Python
Python实现股市信息下载的方法
Jun 15 #Python
给Python入门者的一些编程建议
Jun 15 #Python
You might like
解析php二分法查找数组是否包含某一元素
2013/05/23 PHP
php操作XML、读取数据和写入数据的实现代码
2014/08/15 PHP
Centos PHP 扩展Xchche的安装教程
2016/07/09 PHP
PHP给源代码加密的几种方法汇总(推荐)
2018/02/06 PHP
基于jquery的一个浮动框(扩展性比较好 )
2010/08/27 Javascript
js arguments,jcallee caller用法总结
2013/11/30 Javascript
Javascript中匿名函数的多种调用方式总结
2013/12/06 Javascript
connect中间件session、cookie的使用方法分享
2014/06/17 Javascript
JavaScript设计模式之观察者模式(发布者-订阅者模式)
2014/09/24 Javascript
使用C++为node.js写扩展模块
2015/04/22 Javascript
nodejs爬虫抓取数据之编码问题
2015/07/03 NodeJs
如何实现JavaScript动态加载CSS和JS文件
2020/12/28 Javascript
js获取Html元素的实际宽度高度的方法
2016/05/19 Javascript
基于Javascript实现的不重复ID的生成器
2016/12/25 Javascript
使用Javascript简单计算器
2018/11/17 Javascript
Jquery的autocomplete插件用法及参数讲解
2019/03/12 jQuery
浅谈微信小程序列表埋点曝光指南
2019/10/15 Javascript
基于ajax实现上传图片代码示例解析
2020/12/03 Javascript
使用简单工厂模式来进行Python的设计模式编程
2016/03/01 Python
python机器学习理论与实战(一)K近邻法
2021/01/28 Python
python自动查询12306余票并发送邮箱提醒脚本
2018/05/21 Python
使用python存储网页上的图片实例
2018/05/22 Python
Python pandas.DataFrame调整列顺序及修改index名的方法
2019/06/21 Python
Python+Selenium+phantomjs实现网页模拟登录和截图功能(windows环境)
2019/12/11 Python
详解用Pytest+Allure生成漂亮的HTML图形化测试报告
2020/03/31 Python
html5中canvas学习笔记2-判断浏览器是否支持canvas
2013/01/06 HTML / CSS
瑞典Happy Socks美国官网:购买色彩斑斓的快乐袜子
2016/10/19 全球购物
口腔医学技术应届生求职信
2013/11/09 职场文书
初三班主任寄语大全
2014/04/04 职场文书
贸易跟单员英文求职信
2014/04/19 职场文书
感恩老师的演讲稿
2014/05/06 职场文书
上海世博会志愿者口号
2014/06/17 职场文书
政风行风建设责任书
2014/07/23 职场文书
2015年高一班主任工作总结
2015/05/13 职场文书
关于运动会的广播稿
2015/08/19 职场文书
springboot项目以jar包运行的操作方法
2021/06/30 Java/Android