This is an old revision of the document!
Table of Contents
Html
Layer: Core · Source: lib/core/Html.php:21 (lines 21–584)
final class Html
HTML parsing, filtering and sanitization
This class depends on Tidy which is included in the core since PHP 5.3 Usage: $data = $_POST['body']; $html = new Html(); $data = $html→filter($data);
Docblock Metadata
| Tag | Value |
|---|---|
@category | Core Class |
@author | Eksith Rodrigo <reksith at gmail.com> |
@license | http://opensource.org/licenses/ISC ISC License |
@version | 0.2 |
Inheritance
No parent, interface or trait. This is a root type.
Constants (0)
None.
Properties (3)
| Visibility | Type | Name | Default | Line |
|---|---|---|---|---|
public static | static | $options | array( 'rx_url' ⇒ // URLs over 255 chars can cause problem… | 26 |
private static | static | $tidy | array( // Preserve whitespace inside tags 'add-xml-space' =… | 52 |
private static | static | $whitelist | array( 'p' ⇒ array( 'style', 'class', 'align' ), 'div' ⇒ … | 113 |
Methods (10)
| Visibility | Method | Summary | Line |
|---|---|---|---|
| protected | escapeCode() | Convert content between code blocks into code tags | 181 |
| protected | makeParagraphs() | Convert an unformatted text block to paragraphs | 197 |
| public | filter() | Filters HTML content through whitelist of tags and attributes | 231 |
| protected | cleanAttributeNode() | @SuppressWarnings(PHPMD.ElseExpression) | 302 |
| protected static | linkAttributes() | Modify links to display their domains and add 'nofollow'. | 360 |
| protected | cleanNodes() | Iterate through each tag and add non-whitelisted tags to the | 397 |
| public static | urlFilter() | Returns true if the URL passed value is harmless. | 468 |
| public static | decodeScrub() | Regular expressions don't work well when used for validating HTML. | 506 |
| public static | utfdecode() | UTF-8 compatible URL decoding | 563 |
| public static | entities() | HTML safe character entitites in UTF-8 | 576 |
escapeCode()
protected function escapeCode($val)
lines 181–188 (8)
Convert content between code blocks into code tags
| Parameter | Type | Default | Description |
|---|---|---|---|
$val | (untyped) | required | none |
Returns: (none declared)
makeParagraphs()
protected function makeParagraphs($val)
lines 197–223 (27)
Convert an unformatted text block to paragraphs
| Parameter | Type | Default | Description |
|---|---|---|---|
$val | (untyped) | required | none |
Returns: (none declared)
filter()
public function filter($val)
lines 231–298 (68)
Filters HTML content through whitelist of tags and attributes
| Parameter | Type | Default | Description |
|---|---|---|---|
$val | (untyped) | required | none |
Returns: (none declared)
cleanAttributeNode()
protected function cleanAttributeNode(&$node, &$attr, &$goodAttributes, &$href)
lines 302–353 (52)
@SuppressWarnings(PHPMD.ElseExpression)
| Parameter | Type | Default | Description |
|---|---|---|---|
$node (by reference) | (untyped) | required | none |
$attr (by reference) | (untyped) | required | none |
$goodAttributes (by reference) | (untyped) | required | none |
$href (by reference) | (untyped) | required | none |
Returns: (none declared)
linkAttributes()
protected static function linkAttributes(&$node, $href)
lines 360–387 (28)
Modify links to display their domains and add 'nofollow'.
Also puts the linked domain in the title as well as the file name
| Parameter | Type | Default | Description |
|---|---|---|---|
$node (by reference) | (untyped) | required | none |
$href | (untyped) | required | none |
Returns: (none declared)
cleanNodes()
protected function cleanNodes($node, &$badTags = array())
lines 397–452 (56)
Iterate through each tag and add non-whitelisted tags to the
bad list. Also filter the attributes and remove non-whitelisted ones.
| Parameter | Type | Default | Description |
|---|---|---|---|
$node | (untyped) | required | Current HTML node |
$badTags (by reference) | (untyped) | array() | Cumulative list of tags for deletion |
Returns: (none declared)
urlFilter()
public static function urlFilter($v)
lines 468–495 (28)
Returns true if the URL passed value is harmless.
This regex takes into account Unicode domain names however, it doesn't check for TLD (.com, .net, .mobi, .museum etc…) as that list is too long. The purpose is to ensure your visitors are not harmed by invalid markup, not that they get a functional domain name.
| Parameter | Type | Default | Description |
|---|---|---|---|
$v | (untyped) | required | Raw URL to validate |
Returns: (none declared)
decodeScrub()
public static function decodeScrub($v)
lines 506–554 (49)
Regular expressions don't work well when used for validating HTML.
It really shines when evaluating text so that's what we're doing here
| Parameter | Type | Default | Description |
|---|---|---|---|
$v | (untyped) | required | string Attribute name |
Returns: (none declared)
utfdecode()
public static function utfdecode($v)
lines 563–568 (6)
UTF-8 compatible URL decoding
| Parameter | Type | Default | Description |
|---|---|---|---|
$v | (untyped) | required | none |
Returns: (none declared)
entities()
public static function entities($v)
lines 576–583 (8)
HTML safe character entitites in UTF-8
| Parameter | Type | Default | Description |
|---|---|---|---|
$v | (untyped) | required | none |
Returns: (none declared)
This page is generated from source by 'tools/gendoc'. Edits will be overwritten.
