Sign Up for Free

RunKit +

Try any Node.js package right in your browser

This is a playground to test code. It runs a full Node.js environment and already has all of npm’s 400,000 packages pre-installed, including read-art with all npm packages installed. Try it out:

var readArt = require("read-art")

This service is provided by RunKit and is not affiliated with npm, Inc or the package authors.

read-art v0.5.6

Scrape/Crawl article from any site automatically. Make any web page readable, no matter Chinese or English.

read-art NPM version Build Status js-standard-style

NPM

  1. Readability reference to Arc90's.
  2. Scrape article from any page (automatically).
  3. Make any web page readable, no matter Chinese or English.

快速抓取网页文章标题和内容,适合node.js爬虫使用,服务于ElasticSearch。

Guide

How it works

In my case, the speed of spider is about 1500k documents per day, and the maximize crawling speed is 1.2k /minute, avg 1k /minute, the memory cost are about 200 MB on each spider kernel, and the accuracy is about 90%, the rest 10% can be fixed by customizing Score Rules or Selectors. it's better than any other readability modules.

(4) Server infos:

  • 20M bandwidth of fibre-optical
  • 8 Intel(R) Xeon(R) CPU E5-2650 v2 @ 2.60GHz cpus
  • 32G memory

Metadata

RunKit is a free, in-browser JavaScript dev environment for prototyping Node.js code, with every npm package installed. Sign up to share your code.
Sign Up for Free