Commit Graph
55 Commits
Author SHA1 Message Date
Jan Tojnar 9ed89bde92 Fix PHP-Cs-Fixer changes
1) src/Readability.php (braces, no_unneeded_control_parentheses, single_line_comment_spacing, global_namespace_import, no_unused_imports, phpdoc_align)
   2) src/JSLikeHTMLElement.php (phpdoc_separation)

Switch code blocks to Markdown syntax to work around `phpdoc_separation`, ApiGen uses Markdown these days anyway.
2023-03-31 03:14:00 +02:00
Kevin DecherfandJeremy Benoist 6689f19956 Strip script and style tags through ::clean() method instead of preg_replace
Huge tags can lead to a failure of preg_replace, thus erasing the whole
fetched content.

Fixes https://github.com/wallabag/wallabag/issues/5847

Signed-off-by: Kevin Decherf <kevin@kdecherf.com>
2022-06-13 09:13:23 +02:00
Kevin Decherf 2ab87d7445 Fix isPhrasingContent conditions, text node replacement
It also disables reverting forced paragraph elements as it can break
layouts or corrupt content.

Signed-off-by: Kevin Decherf <kevin@kdecherf.com>
2022-02-15 20:53:28 +01:00
Jeremy Benoist c2a1639b34 Add Rector 2022-02-04 12:13:37 +01:00
Kevin Decherf a44c4e5482 Add routine to remove invisible nodes
Readability was previously removing (was trying to actually, see next
section) invisible nodes using a pattern from `unlikelyCandidates`. This
was quite hacky and was removed during a backport of logics from
mozilla/readability. There is still a need to remove them so here we
are. We still use a pattern but specifically against the style
attribute. We also remove nodes with the attribute `hidden`.

The clean feature of tidy actually replaces inline style attributes
with css classes thus preventing readability to detect invisible nodes,
see https://github.com/htacg/tidy-html5/blob/5.6.0/src/clean.c#L1488
We therefore set clean configuration to false.

Signed-off-by: Kevin Decherf <kevin@kdecherf.com>
2022-02-04 01:23:11 +01:00
Kevin Decherf b580cf216d Backport some logics from mozilla/readability
This change backports several things from mozilla/readability:

- Add child score to all ancestors instead of the first parent only
- Check 5 top candidates and try to find alternative candidates within
  ancestors, this can help to find a better parent and grab more content
- Reduce patterns from `unlikelyCandidates` to the one used by Mozilla
  as ours tend to remove useful nodes
- Score headers (h2 to h6) by default in addition to div, p, td and
  section

Signed-off-by: Kevin Decherf <kevin@kdecherf.com>
2022-02-04 01:05:13 +01:00
Jeremy Benoist c4bba53dbe Remove Scrutinizer 2022-02-02 12:52:12 +01:00
Jeremy Benoist 66215a6c80 Require PHP >= 7.2
- remove test on Composer v1
- remove deprecated function
- move `loadHtml()` into `init()` instead of `__construct`

Kinda prepare 2.0 version :)
2022-02-02 12:44:24 +01:00
Jérémy BenoistandGitHub fabf096ce6 Fix deprecated message
> Method "Psr\Log\LoggerAwareInterface::setLogger()" might add "void" as a native return type declaration in the future. Do the same in implementation "Readability\Readability" now to avoid errors or add an explicit @return annotation to suppress this message.
2021-11-29 20:50:10 +01:00
Kevin Decherf eb72a315c4 Clean empty figure tags without ending
See 'Tag omission' https://developer.mozilla.org/en-US/docs/Web/HTML/Element/figure

Signed-off-by: Kevin Decherf <kevin@kdecherf.com>
2021-10-29 15:36:19 +02:00
Jérémy BenoistandGitHub 9a490fac07 Merge pull request #52 from nicofrand/master
Skip empty (empty innerHTML) nodes when grabbing article
2021-03-09 11:14:29 +01:00
Jeremy Benoist ea1368fac0 Body can be wiped without tidy
Re-create it in that case.

Also run CS-Fixer.
2021-03-08 11:59:24 +01:00
Jan TojnarandGitHub 7cea79c23a readability: stop tidy from wrapping noscript text
HTML 4.01 Strict only allows block-level elements within noscript, form and
blockquote. The `enclose-block-text` option fixes the instances when those
elements contain inline elements or text by wrapping the children in paragraphs.

HTML 5 has looser content model and allows noscript elements basically anywhere,
including paragraphs, making the noscript elements inherit the parent element’s
content model. This means that tidy will produce invalid HTML nesting paragraphs
for `p > noscript > text`, a structure that would be invalid on two counts
in HTML 4 Strict profile but is completely valid in HTML 5.

Popular WordPress image lazy-loading code produces precisely that structure
so tidy “corrects” it to invalid code. In a proper HTML parser, the produced
code would force close the outer paragraph, making the noscript element
its sibling instead of a child. The only reason this does not break Graby’s code
for stripping the lazy-loading HTML is that libxml2 contains a bug
counteracting this:

https://gitlab.gnome.org/GNOME/libxml2/-/issues/205

Since all three elements allow flow content in HTML 5, it does not make much
sense to enable this option any more. The only possible issues that could occur
is producing HTML code not conforming to 4.01 Strict but that was never guaranteed,
as our example shows, and having blockquotes contain text nodes not wrapped
in paragraphs, which might be expected by some ancient stylesheets
but that is only minor and easily fixable visual backwards incompatibility.
2020-11-14 22:13:53 +01:00
Jeremy Benoist 6a8ecf232f Use a new deps for HTML5 parser
`electrolinux/php-html5lib` was quite old and incompatible with the upcoming Composer 2.0.
Jumping to `masterminds/html5` for the same result. Also the lib is maintained.

Also:
- keep README in vendors
- use new Scrutinizer engine
- test with lower deps
- remove php-coveralls dev deps and download the phar during the CI build
2020-06-08 07:04:00 +02:00
Jeremy Benoist b1acc9ed73 Fix PHPStan (again)
Also cleanup
2019-11-19 14:09:29 +01:00
Jeremy Benoist 11d2946904 Add openload.co to media detection 2019-06-25 16:54:38 +02:00
nicofrand ff78c63e6d Skip empty (empty innerHTML) nodes when grabbing article 2019-05-25 16:12:52 +02:00
Jeremy Benoist bb65caf864 Fix “A non well formed numeric value encountered” 2019-05-11 21:58:11 +02:00
Simounet 2e20f76195 \bout removed from negative content 2019-04-19 12:12:41 +02:00
Jeremy Benoist 74d9cc605a Enable PHPStan 2019-02-07 15:51:31 +01:00
Jeremy Benoist 2dce2879bf Update fixer rules
Following graby, wallabag, etc.
2019-02-04 11:21:34 +01:00
Kevin DecherfandJeremy Benoist 26c881d864 tidy: use tidy_repair_string instead of tidy_parse_string+tidy_clean_repair
A change released in tidy 5.6.0 breaks php-tidy when using
tidy_parse_string+tidy_clean_repair and wrap=0, incorrectly wrapping
every single word. Also it seems that $tidy->value should not be used to
retrieve the repaired html as far as it is undocumented and for internal
use.

We replace the call with tidy_repair_string which directly returns the
repaired string.

Relates to https://github.com/htacg/tidy-html5/issues/673
Relates to https://bugs.php.net/bug.php?id=75947

Tests pass.

Signed-off-by: Kevin Decherf <kevin@kdecherf.com>
2019-02-04 11:08:18 +01:00
Simounet 422c74f29c Giphy added to allowed medias 2018-11-20 19:27:41 +01:00
Simounet 63cd304dba Media class added to positive candidates
Fix Mediapart images.
2018-06-05 13:26:39 +02:00
Kevin Decherf 4c68cc9f09 Keep elements with 'footnote' as possible candidates
Should fix https://github.com/wallabag/wallabag/issues/3100

Signed-off-by: Kevin Decherf <kevin@kdecherf.com>
2017-11-01 16:45:07 +01:00
Jeremy Benoist 613a63c062 CS 2017-06-30 16:42:29 +02:00
Jeremy Benoist 05089bbd03 Add missing HTML5 class 2017-06-30 16:32:37 +02:00
Jeremy Benoist f2a43b476c Avoid PHP Warning
This isn't the best solution but the previous one using `@` wasn't really better.
Appending a string into a fragment might generate some warning if the string contains bad entity.
For example `&plus;`.
2017-05-19 15:37:46 +02:00
Jeremy Benoist 8b1c3f147d Don't be to hard on 'links' attribute 2017-02-02 15:57:15 +01:00
Jeremy Benoist ff754b80bd Avoid childnode becoming null to generate a warning 2017-01-10 10:58:12 +01:00
Jeremy Benoist d97bece7c5 Don’t be too agressive
Some links got a “tooltip-link” and shouldn’t be removed by php-readability because they are usefull to the content
2016-10-20 23:40:52 +02:00
Jeremy Benoist 3de4e918b4 Convert header & section to p
And took `pre` element in score
2016-10-02 14:55:52 +02:00
Jeremy Benoist 5182d6cb11 “info” is too agressive in unlikelyCandidates
Some contents have a `infocontent` node (ot sth different) and they are real content.
Using only `info` as regex is too agressive and remove legitimate content.
Matching the whole word `info` (or `infos`) should be a better choice
2016-10-02 14:49:43 +02:00
Jeremy Benoist 2ef400bf73 Enable php-cs-fixer 2016-06-23 07:28:10 +02:00
Jeremy Benoist 00f622e9b7 Revert BC changes
- avoid method signature update
- revert moving logic out of the constructor
2016-03-01 15:07:32 +01:00
Jeremy Benoist c756ec067e Fix tests
`getInnerText` might receive a null DOMElement if the xpath or query return no element.
2016-02-29 13:09:16 +01:00
Jeremy Benoist 8ab7d76cd5 Use Monolog instead of custom solution
Remove that ugly `openlog` & `syslog`
2016-02-29 13:01:40 +01:00
Jeremy Benoist 149a333b40 Remove addPreFilter
Pre filters are used in the __construct so adding more pre filters once the object is instantiated is useless.
2016-02-29 12:10:56 +01:00
Jeremy Benoist 209c404d7b Fix instanceof DOMElement
We previously checked `instanceof DOMElement` which was wrong since we
are in the namespace class, the class `Readability\DOMElement` does not
exists.
2016-02-29 11:53:03 +01:00
Jeremy Benoist 2951936e00 CS & PHPDoc 2016-02-29 11:00:45 +01:00
Jeremy Benoist 850ade16b6 Cleanup 2016-02-29 10:22:07 +01:00
Jeremy Benoist dc590542f0 Avoid adding id that might already exists
We append a new node when it isn't a `div` or `p` (like when it's an `article`) with the same id which generate a DOM error "DOMElement::setAttribute(): ID blabla already defined".
2016-02-29 10:21:52 +01:00
Jeremy Benoist 111cb08034 Improve negative element
- add unlikelyCandidates: head
- add negative: recommend
2015-11-09 19:44:11 +01:00
Jeremy Benoist f71c3a4196 Do not remove html tag attributes
They might contains useful information (at least language)
2015-09-23 21:09:38 +02:00
Jeremy Benoist b77876b30a Do not remove nofollow links
Most the time, they can be usefull.
At least, it'll be a link to something unrelated. But we won't lose a link inside the content.

Also, adding some extra space.
2015-09-22 19:25:57 +02:00
Jeremy Benoist 175196d6c2 Avoid error with &nbsp;
Fix #5
2015-09-18 19:10:48 +02:00
Jeremy Benoist 2b5af601d5 Do not format output to avoid breaking apps
It'll require to jump to 2.0.0 and I think it's to soon
2015-09-15 22:25:08 +02:00
Jeremy Benoist d01eb2ac1e Use class instead of id to avoid error
It generates error like `ID XXX already defined`
2015-09-14 21:49:40 +02:00
Jeremy Benoist c5a4a490e1 CS 2015-08-24 11:10:54 +02:00
Jeremy Benoist c67189248e Backport changes from wallabag
https://github.com/wallabag/php-readability/commit/e9e4ff87f8fc56d406ccdd5a9a7f1d3d6af07e79
2015-08-24 11:09:38 +02:00