Scriptlog Docs

Scriptlog Documentation

Code reference for the Scriptlog codebase

User Tools

Site Tools


scriptlog:lib:core:html

Html

Layer: Core · Source: lib/core/Html.php:21 (lines 21–584)


final class Html

HTML parsing, filtering and sanitization

This class depends on Tidy which is included in the core since PHP 5.3 Usage: $data = $_POST['body']; $html = new Html(); $data = $html→filter($data);

Docblock Metadata

Tag Value
@category Core Class
@author Eksith Rodrigo <reksith at gmail.com>
@license http://opensource.org/licenses/ISC ISC License
@version 0.2

Inheritance

No parent, interface or trait. This is a root type.

Constants (0)

None.

Properties (3)

Visibility Type Name Default Line
public static static $options array( 'rx_url' ⇒ // URLs over 255 chars can cause problem… 26
private static static $tidy array( // Preserve whitespace inside tags 'add-xml-space' =… 52
private static static $whitelist array( 'p' ⇒ array( 'style', 'class', 'align' ), 'div' ⇒ … 113

Methods (10)

Visibility Method Summary Line
protected escapeCode() Convert content between code blocks into code tags 181
protected makeParagraphs() Convert an unformatted text block to paragraphs 197
public filter() Filters HTML content through whitelist of tags and attributes 231
protected cleanAttributeNode() 302
protected static linkAttributes() Modify links to display their domains and add 'nofollow'. 360
protected cleanNodes() Iterate through each tag and add non-whitelisted tags to the 397
public static urlFilter() Returns true if the URL passed value is harmless. 468
public static decodeScrub() Regular expressions don't work well when used for validating HTML. 506
public static utfdecode() UTF-8 compatible URL decoding 563
public static entities() HTML safe character entitites in UTF-8 576

escapeCode()

protected function escapeCode($val)

lines 181–188 (8)

Convert content between code blocks into code tags

Parameter Type Default Description
$val (untyped) required none

Returns: (none declared)

makeParagraphs()

protected function makeParagraphs($val)

lines 197–223 (27)

Convert an unformatted text block to paragraphs

Parameter Type Default Description
$val (untyped) required none

Returns: (none declared)

filter()

public function filter($val)

lines 231–298 (68)

Filters HTML content through whitelist of tags and attributes

Parameter Type Default Description
$val (untyped) required none

Returns: (none declared)

cleanAttributeNode()

protected function cleanAttributeNode(&amp;$node, &amp;$attr, &amp;$goodAttributes, &amp;$href)

lines 302–353 (52)

Parameter Type Default Description
$node (by reference) (untyped) required none
$attr (by reference) (untyped) required none
$goodAttributes (by reference) (untyped) required none
$href (by reference) (untyped) required none

Returns: (none declared)

linkAttributes()

protected static function linkAttributes(&amp;$node, $href)

lines 360–387 (28)

Modify links to display their domains and add 'nofollow'.

Also puts the linked domain in the title as well as the file name

Parameter Type Default Description
$node (by reference) (untyped) required none
$href (untyped) required none

Returns: (none declared)

cleanNodes()

protected function cleanNodes($node, &amp;$badTags = array())

lines 397–452 (56)

Iterate through each tag and add non-whitelisted tags to the

bad list. Also filter the attributes and remove non-whitelisted ones.

Parameter Type Default Description
$node (untyped) required Current HTML node
$badTags (by reference) (untyped) array() Cumulative list of tags for deletion

Returns: (none declared)

urlFilter()

public static function urlFilter($v)

lines 468–495 (28)

Returns true if the URL passed value is harmless.

This regex takes into account Unicode domain names however, it doesn't check for TLD (.com, .net, .mobi, .museum etc…) as that list is too long. The purpose is to ensure your visitors are not harmed by invalid markup, not that they get a functional domain name.

Parameter Type Default Description
$v (untyped) required Raw URL to validate

Returns: (none declared)

decodeScrub()

public static function decodeScrub($v)

lines 506–554 (49)

Regular expressions don't work well when used for validating HTML.

It really shines when evaluating text so that's what we're doing here

Parameter Type Default Description
$v (untyped) required string Attribute name

Returns: (none declared)

utfdecode()

public static function utfdecode($v)

lines 563–568 (6)

UTF-8 compatible URL decoding

Parameter Type Default Description
$v (untyped) required none

Returns: (none declared)

entities()

public static function entities($v)

lines 576–583 (8)

HTML safe character entitites in UTF-8

Parameter Type Default Description
$v (untyped) required none

Returns: (none declared)


This page is generated from source by 'tools/gendoc'. Edits will be overwritten.

scriptlog/lib/core/html.txt · Last modified: by admin