Scrapy

Scrapy([ˈskreɪpaɪ] SKRAY-peye)はPythonで開発されたフリーでオープンソースのクロールフレームワーク。元々はウェブスク…

Scrapy
Scrapy
開発元 Scrapinghub, Ltd.
初版 2008年6月26日 (2008-06-26)
最新版 2.18.0[1] ウィキデータを編集 - 2026年8月20日 (4日前) [±]
リポジトリ ウィキデータを編集
プログラミング
言語
Python
対応OS Windows, macOS, Linux
種別 Web crawler
ライセンス BSD License
公式サイト scrapy.org ウィキデータを編集
テンプレートを表示

Scrapy[ˈskrp] SKRAY-peye)はPythonで開発されたフリーでオープンソースクロールフレームワーク。元々はウェブスクレイピング用に設計されたが、 APIを使用したデータの抽出や、汎用のクローラーとしても使用できる[2]。現在、ウェブスクレイピングの開発およびサービス会社であるScrapinghub Ltd.で管理されている。 Scrapyプロジェクトアーキテクチャは、「スパイダー」を中心に構築されている。DjangoなどのフレームワークをDRY[3]他の精神を踏襲し、開発者がコードを再利用できるようにしている。 さらに、サイトの動作に関する想定をテストするために開発者が使用できるWebクロールシェルを提供する[4]。 Scrapyを使用している有名な会社と製品には、Lyst[5][6]、Parse.ly[7]、Sayone Technologies[8]Sciences Po Medialab[9]、Data.gov.ukの世界政府データサイト[10]がある[11]

Scrapyは、ロンドンを拠点とするアグリゲーターおよびEC会社のMydecoで開発がスタートした。Mydecoは、MydecoおよびInsophia(ウルグアイのモンテビデオに拠点を置くWebコンサルティング会社)の従業員によって開発および管理されている。 最初の公開リリースはBSDライセンスに基づく2008年8月で、マイルストーン1.0のリリースは2015年6月に行われた。 2011年に、Scrapinghubが新しい公式メンテナになった[12][13]

出典

  1. ^ Release 2.18.0” (2026年8月20日). 2026年8月21日閲覧。
  2. ^ Scrapy at a glance.
  3. ^ Frequently Asked Questions”. 2015年7月28日閲覧。
  4. ^ Scrapy shell”. 2015年7月28日閲覧。
  5. ^ Bell. “Scalable Scraping Using Machine Learning”. 2015年7月28日閲覧。
  6. ^ Scrapy | Companies using Scrapy
  7. ^ Montalenti (2012年10月27日). “Web Crawling & Metadata Extraction in Python”. 2020年8月4日閲覧。
  8. ^ Scrapy Companies”. Scrapy website. 2020年8月4日閲覧。
  9. ^ Hyphe v0.0.0: the first release of our new webcrawler is out!
  10. ^ Ben Firshman [@bfirsh] (21 January 2010). “World Govt Data site uses Django, Solr, Haystack, Scrapy and other exciting buzzwords http://bit.ly/5jU3La #opendata #datastore”. X(旧Twitter)より2020年8月4日閲覧.
  11. ^ [1]
  12. ^ Pablo Hoffman (2013). List of the primary authors & contributors. https://github.com/scrapy/scrapy/blob/master/AUTHORS 2013年11月18日閲覧。 
  13. ^ Interview Scraping Hub.

外部リンク

Content Disclaimer

Informasi ini disarikan dari Wikipedia dan disajikan kembali untuk tujuan edukasi. Konten tersedia di bawah lisensi CC BY-SA 3.0. Kami tidak bertanggung jawab atas ketidakakuratan data yang bersumber dari kontribusi publik tersebut.

  1. The information displayed on this website is sourced in part or in whole from Wikipedia and has been adapted for the purpose of restating it. We strive to provide accurate and relevant information, however:
  2. There is no guarantee of absolute accuracy. Wikipedia is an open, collaborative project that can be edited by anyone, so information is subject to change.
  3. It is not intended to constitute professional advice. The content displayed is for informational and educational purposes only. For important decisions (e.g., medical, legal, or financial), please consult a professional.
  4. Content copyright. Wikipedia is licensed under the Creative Commons Attribution-ShareAlike License (CC BY-SA). This means that content may be reused with appropriate attribution and shared under a similar license.
  5. Responsible use. Any risk arising from the use of information from this website is entirely the responsibility of the user.