首页 > 其他 > 详细

[scrapy] scrapy 使用goose作为正文提取

时间:2015-08-25 19:28:19      阅读:142      评论:0      收藏:0      [点我收藏+]
import scrapy
from goose import Goose

class Article(scrapy.Item):
    title = scrapy.Field()
    text = scrapy.Field()

class MyGooseSpider(scrapy.Spider):
    name = ‘goose‘
    start_urls = [
        ‘http://blog.scrapinghub.com/2014/06/18/extracting-schema-org-microdata-using-scrapy-selectors-and-xpath/‘,
        ‘http://blog.scrapinghub.com/2014/07/17/xpath-tips-from-the-web-scraping-trenches/‘,
    ]

    def parse(self, response):
        article = Goose().extract(raw_html=response.body)
        yield Article(title=article.title, text=article.cleaned_text)

转自:http://stackoverflow.com/questions/26940002/can-i-use-scrapy-with-goose

[scrapy] scrapy 使用goose作为正文提取

原文:http://www.cnblogs.com/bushe/p/4757981.html

(0)
(0)
   
举报
评论 一句话评论(0
关于我们 - 联系我们 - 留言反馈 - 联系我们:wmxa8@hotmail.com
© 2014 bubuko.com 版权所有
打开技术之扣,分享程序人生!