Python实现简单HTML表格解析的方法


Posted in Python onJune 15, 2015

本文实例讲述了Python实现简单HTML表格解析的方法。分享给大家供大家参考。具体分析如下:

这里依赖libxml2dom,确保首先安装!导入到你的脚步并调用parse_tables() 函数。

1. source = a string containing the source code you can pass in just the table or the entire page code

2. headers = a list of ints OR a list of strings
If the headers are ints this is for tables with no header, just list the 0 based index of the rows in which you want to extract data.
If the headers are strings this is for tables with header columns (with the tags) it will pull the information from the specified columns

3. The 0 based index of the table in the source code. If there are multiple tables and the table you want to parse is the third table in the code then pass in the number 2 here

It will return a list of lists. each inner list will contain the parsed information.

具体代码如下:

#The goal of table parser is to get specific information from specific
#columns in a table.
#Input: source code from a typical website
#Arguments: a list of headers the user wants to return
#Output: A list of lists of the data in each row
import libxml2dom
def parse_tables(source, headers, table_index):
  """parse_tables(string source, list headers, table_index)
    headers may be a list of strings if the table has headers defined or
    headers may be a list of ints if no headers defined this will get data
    from the rows index.
    This method returns a list of lists
    """
  #Determine if the headers list is strings or ints and make sure they
  #are all the same type
  j = 0
  print 'Printing headers: ',headers
  #route to the correct function
  #if the header type is int
  if type(headers[0]) == type(1):
    #run no_header function
    return no_header(source, headers, table_index)
  #if the header type is string
  elif type(headers[0]) == type('a'):
    #run the header_given function
    return header_given(source, headers, table_index)
  else:
    #return none if the headers aren't correct
    return None
#This function takes in the source code of the whole page a string list of
#headers and the index number of the table on the page. It returns a list of
#lists with the scraped information
def header_given(source, headers, table_index):
  #initiate a list to hole the return list
  return_list = []
  #initiate a list to hold the index numbers of the data in the rows
  header_index = []
  #get a document object out of the source code
  doc = libxml2dom.parseString(source,html=1)
  #get the tables from the document
  tables = doc.getElementsByTagName('table')
  try:
    #try to get focue on the desired table
    main_table = tables[table_index]
  except:
    #if the table doesn't exits then return an error
    return ['The table index was not found']
  #get a list of headers in the table
  table_headers = main_table.getElementsByTagName('th')
  #need a sentry value for the header loop
  loop_sentry = 0
  #loop through each header looking for matches
  for header in table_headers:
    #if the header is in the desired headers list 
    if header.textContent in headers:
      #add it to the header_index
      header_index.append(loop_sentry)
    #add one to the loop_sentry
    loop_sentry+=1
  #get the rows from the table
  rows = main_table.getElementsByTagName('tr')
  #sentry value detecting if the first row is being viewed
  row_sentry = 0
  #loop through the rows in the table, skipping the first row
  for row in rows:
    #if row_sentry is 0 this is our first row
    if row_sentry == 0:
      #make the row_sentry not 0
      row_sentry = 1337
      continue
    #get all cells from the current row
    cells = row.getElementsByTagName('td')
    #initiate a list to append into the return_list
    cell_list = []
    #iterate through all of the header index's
    for i in header_index:
      #append the cells text content to the cell_list
      cell_list.append(cells[i].textContent)
    #append the cell_list to the return_list
    return_list.append(cell_list)
  #return the return_list
  return return_list
#This function takes in the source code of the whole page an int list of
#headers indicating the index number of the needed item and the index number
#of the table on the page. It returns a list of lists with the scraped info
def no_header(source, headers, table_index):
  #initiate a list to hold the return list
  return_list = []
  #get a document object out of the source code
  doc = libxml2dom.parseString(source, html=1)
  #get the tables from document
  tables = doc.getElementsByTagName('table')
  try:
    #Try to get focus on the desired table
    main_table = tables[table_index]
  except:
    #if the table doesn't exits then return an error
    return ['The table index was not found']
  #get all of the rows out of the main_table
  rows = main_table.getElementsByTagName('tr')
  #loop through each row
  for row in rows:
    #get all cells from the current row
    cells = row.getElementsByTagName('td')
    #initiate a list to append into the return_list
    cell_list = []
    #loop through the list of desired headers
    for i in headers:
      try:
        #try to add text from the cell into the cell_list
        cell_list.append(cells[i].textContent)
      except:
        #if there is an error usually an index error just continue
        continue
    #append the data scraped into the return_list    
    return_list.append(cell_list)
  #return the return list
  return return_list

希望本文所述对大家的Python程序设计有所帮助。

Python 相关文章推荐
Python编码类型转换方法详解
Jul 01 Python
为什么入门大数据选择Python而不是Java?
Mar 07 Python
Python callable()函数用法实例分析
Mar 17 Python
pandas 将list切分后存入DataFrame中的实例
Jul 03 Python
Python调用服务接口的实例
Jan 03 Python
Python中的异常处理try/except/finally/raise用法分析
Feb 28 Python
简单了解python关系(比较)运算符
Jul 08 Python
python PIL和CV对 图片的读取,显示,裁剪,保存实现方法
Aug 07 Python
Python学习笔记之For循环用法详解
Aug 14 Python
Python笔记之工厂模式
Nov 20 Python
Numpy 理解ndarray对象的示例代码
Apr 03 Python
解决启动django,浏览器显示“服务器拒绝访问”的问题
May 13 Python
Python判断Abundant Number的方法
Jun 15 #Python
Python计算一个文件里字数的方法
Jun 15 #Python
Python素数检测实例分析
Jun 15 #Python
Python计算三维矢量幅度的方法
Jun 15 #Python
Python栈类实例分析
Jun 15 #Python
Python实现股市信息下载的方法
Jun 15 #Python
给Python入门者的一些编程建议
Jun 15 #Python
You might like
PHP、Python和Javascript的装饰器模式对比
2015/02/03 PHP
php 指定范围内多个随机数代码实例
2016/07/18 PHP
PHP获取路径和目录的方法总结【必看篇】
2017/03/04 PHP
PHP get_html_translation_table()函数用法讲解
2019/02/16 PHP
php实现的支付宝网页支付功能示例【基于TP5框架】
2019/09/16 PHP
javascript里的条件判断
2007/02/27 Javascript
Javascript和Ajax中文乱码吐血版解决方案
2009/12/21 Javascript
js 操作select和option常用代码整理
2012/12/13 Javascript
node.js中的fs.lchown方法使用说明
2014/12/16 Javascript
jQuery使用元素属性attr赋值详解
2015/02/27 Javascript
jQuery使用toggleClass方法动态添加删除Class样式的方法
2015/03/26 Javascript
JQuery中Text方法用法实例分析
2015/05/18 Javascript
JavaScript实现点击按钮切换网页背景色的方法
2015/10/17 Javascript
微信小程序 下拉菜单简单实例
2017/04/13 Javascript
spring+angular实现导出excel的实现代码
2019/02/27 Javascript
jQuery实现form表单基于ajax无刷新提交方法实例代码
2019/11/04 jQuery
js实现时间日期校验
2020/05/26 Javascript
解决vue init webpack 下载依赖卡住不动的问题
2020/11/09 Javascript
python结合API实现即时天气信息
2016/01/19 Python
Python中查看文件名和文件路径
2017/03/31 Python
python web基础之加载静态文件实例
2018/03/20 Python
python交换两个变量的值方法
2019/01/12 Python
python主线程与子线程的结束顺序实例解析
2019/12/17 Python
python对Excel的读取的示例代码
2020/02/14 Python
python模拟点击网页按钮实现方法
2020/02/25 Python
python 写一个水果忍者游戏
2021/01/13 Python
VIVOBAREFOOT赤脚鞋:让您的脚做自然的事情
2017/06/01 全球购物
英国假睫毛购买网站:FalseEyelashes.co.uk
2018/05/23 全球购物
英国蛋糕装饰用品一站式商店:Craft Company
2019/03/18 全球购物
门卫班长岗位职责
2013/12/15 职场文书
新护士岗前培训制度
2014/02/02 职场文书
保护环境倡议书500字
2014/05/19 职场文书
网吧七夕活动策划方案
2014/08/31 职场文书
街道党风廉政建设调研报告
2015/01/01 职场文书
刑事附带民事起诉状
2015/05/19 职场文书
SpringBoot集成MongoDB实现文件上传的步骤
2022/04/18 MongoDB