Table of Contents

Class HtmlUtility

Namespace
Songhay.Xml
Assembly
SonghayCore.dll

Static members for HTML text processing.

public static class HtmlUtility
Inheritance
HtmlUtility
Inherited Members

Methods

ConvertToHtml(string?)

Returns a string of marked up text compatible with browsers that do not support XHTML (loosely towards HTML 4.x W3C standard).

public static string? ConvertToHtml(string? input)

Parameters

input string

A string of markup.

Returns

string

ConvertToXml(string?)

Attempts to convert HTML to well-formed XML with a conventional set of Regex methods.

public static string? ConvertToXml(string? html)

Parameters

html string

An HTML string.

Returns

string

Remarks

This method is simpler than converting to XHTML.

This member uses the following Regex methods:

FormatXhtmlElements(string?)

Formats the specified fragment with a conventional set of Regex methods.

public static string? FormatXhtmlElements(string? xmlFragment)

Parameters

xmlFragment string

A well-formed string of XML.

Returns

string

Remarks

This member uses the following Regex method:

GetBooleanAttributes()

Returns all known HTML Boolean attributes

public static string[] GetBooleanAttributes()

Returns

string[]

Remarks

GetHtmlBooleanAttributesPattern()

Regex pattern getter

public static string GetHtmlBooleanAttributesPattern()

Returns

string

GetInnerXml(string?, string?, string, byte)

Returns the …inner… fragment of XML from the specified unique element.

public static string? GetInnerXml(string? xmlFragment, string? elementName, string newLine = "\r\n", byte numberOfChars = 4)

Parameters

xmlFragment string

A well-formed string of XML.

elementName string

The local name of the element in the XML string.

newLine string

the conventional NewLine characters of the lines

numberOfChars byte

the number of white space characters to remove

Returns

string

MatchHtmlAttributesWithoutQuotes()

[GeneratedRegex("\\s([^'\\\"\\s]+)(\\s*=\\s*)([^'\\\"\\s]+)\\s", RegexOptions.IgnoreCase)]
public static Regex MatchHtmlAttributesWithoutQuotes()

Returns

Regex

Remarks

Pattern:

\\s([^'\\"\\s]+)(\\s*=\\s*)([^'\\"\\s]+)\\s

Options:<br />
<pre><code class="lang-csharp">RegexOptions.IgnoreCase</code></pre><br />
Explanation:<br />
<pre><code class="lang-csharp">○ Match a whitespace character.

○ 1st capture group. ○ Match a character in the set [^"'\s] greedily at least once. ○ 2nd capture group. ○ Match a whitespace character atomically any number of times. ○ Match '='. ○ Match a whitespace character greedily any number of times. ○ 3rd capture group. ○ Match a character in the set [^"'\s] greedily at least once. ○ Match a whitespace character.

MatchHtmlClosingTagCharacters()

[GeneratedRegex("\\<\\s*/")]
public static Regex MatchHtmlClosingTagCharacters()

Returns

Regex

Remarks

Pattern:

\\<\\s*/

Explanation:<br />
<pre><code class="lang-csharp">○ Match '<'.

○ Match a whitespace character atomically any number of times. ○ Match '/'.

MatchHtmlElementsThatShouldNotBeMinimized()

[GeneratedRegex("<(a|iframe|td|th|script)\\s+([^>]*)(/>)", RegexOptions.IgnoreCase)]
public static Regex MatchHtmlElementsThatShouldNotBeMinimized()

Returns

Regex

Remarks

Pattern:

<(a|iframe|td|th|script)\\s+([^>]*)(/>)

Options:<br />
<pre><code class="lang-csharp">RegexOptions.IgnoreCase</code></pre><br />
Explanation:<br />
<pre><code class="lang-csharp">○ Match '<'.

○ 1st capture group. ○ Match with 5 alternative expressions. ○ Match a character in the set [Aa]. ○ Match a sequence of expressions. ○ Match a character in the set [Ii]. ○ Match a character in the set [Ff]. ○ Match a character in the set [Rr]. ○ Match a character in the set [Aa]. ○ Match a character in the set [Mm]. ○ Match a character in the set [Ee]. ○ Match a sequence of expressions. ○ Match a character in the set [Tt]. ○ Match a character in the set [Dd]. ○ Match a sequence of expressions. ○ Match a character in the set [Tt]. ○ Match a character in the set [Hh]. ○ Match a sequence of expressions. ○ Match a character in the set [Ss]. ○ Match a character in the set [Cc]. ○ Match a character in the set [Rr]. ○ Match a character in the set [Ii]. ○ Match a character in the set [Pp]. ○ Match a character in the set [Tt]. ○ Match a whitespace character greedily at least once. ○ 2nd capture group. ○ Match a character other than '>' greedily any number of times. ○ 3rd capture group. ○ Match the string "/>".

MatchHtmlHrefOrSrcAttributes()

[GeneratedRegex("(href|src)\\s*=\\s*['\"][^\"]+['\"]", RegexOptions.IgnoreCase)]
public static Regex MatchHtmlHrefOrSrcAttributes()

Returns

Regex

Remarks

Pattern:

(href|src)\\s*=\\s*['"][^"]+['"]

Options:<br />
<pre><code class="lang-csharp">RegexOptions.IgnoreCase</code></pre><br />
Explanation:<br />
<pre><code class="lang-csharp">○ 1st capture group.
○ Match with 2 alternative expressions.
    ○ Match a sequence of expressions.
        ○ Match a character in the set [Hh].
        ○ Match a character in the set [Rr].
        ○ Match a character in the set [Ee].
        ○ Match a character in the set [Ff].
    ○ Match a sequence of expressions.
        ○ Match a character in the set [Ss].
        ○ Match a character in the set [Rr].
        ○ Match a character in the set [Cc].

○ Match a whitespace character atomically any number of times. ○ Match '='. ○ Match a whitespace character atomically any number of times. ○ Match a character in the set ["']. ○ Match a character other than '"' greedily at least once. ○ Match a character in the set ["'].

MatchHtmlStartTags()

[GeneratedRegex("<[^>\\/]+>")]
public static Regex MatchHtmlStartTags()

Returns

Regex

Remarks

Pattern:

<[^>\\/]+>

Explanation:<br />
<pre><code class="lang-csharp">○ Match '<'.

○ Match a character in the set [^/>] atomically at least once. ○ Match '>'.

MatchHtmlTagContents()

[GeneratedRegex("<[^/][^>]*>")]
public static Regex MatchHtmlTagContents()

Returns

Regex

Remarks

Pattern:

<[^/][^>]*>

Explanation:<br />
<pre><code class="lang-csharp">○ Match '<'.

○ Match any character other than '/'. ○ Match a character other than '>' atomically any number of times. ○ Match '>'.

MatchHtmlTagWithAnyAttributes()

[GeneratedRegex("<html [^>]*>", RegexOptions.IgnoreCase)]
public static Regex MatchHtmlTagWithAnyAttributes()

Returns

Regex

Remarks

Pattern:

<html [^>]*>

Options:<br />
<pre><code class="lang-csharp">RegexOptions.IgnoreCase</code></pre><br />
Explanation:<br />
<pre><code class="lang-csharp">○ Match '<'.

○ Match a character in the set [Hh]. ○ Match a character in the set [Tt]. ○ Match a character in the set [Mm]. ○ Match a character in the set [Ll]. ○ Match ' '. ○ Match a character other than '>' atomically any number of times. ○ Match '>'.

MatchXhtmlAttribute()

[GeneratedRegex("\\s+(.+)\\s*=\\s*\"\\1\"")]
public static Regex MatchXhtmlAttribute()

Returns

Regex

Remarks

Pattern:

\\s+(.+)\\s*=\\s*"\\1"

Explanation:<br />
<pre><code class="lang-csharp">○ Match a whitespace character greedily at least once.

○ 1st capture group. ○ Match a character other than '\n' greedily at least once. ○ Match a whitespace character atomically any number of times. ○ Match '='. ○ Match a whitespace character atomically any number of times. ○ Match '"'. ○ Match the same text as matched by the 1st capture group. ○ Match '"'.

MatchXhtmlEndTagsThatShouldBeMinimized()

[GeneratedRegex("\\s*</(base|isindex|link|meta)\\s*>", RegexOptions.IgnoreCase)]
public static Regex MatchXhtmlEndTagsThatShouldBeMinimized()

Returns

Regex

Remarks

Pattern:

\\s*</(base|isindex|link|meta)\\s*>

Options:<br />
<pre><code class="lang-csharp">RegexOptions.IgnoreCase</code></pre><br />
Explanation:<br />
<pre><code class="lang-csharp">○ Match a whitespace character atomically any number of times.

○ Match the string "</". ○ 1st capture group. ○ Match with 4 alternative expressions. ○ Match a sequence of expressions. ○ Match a character in the set [Bb]. ○ Match a character in the set [Aa]. ○ Match a character in the set [Ss]. ○ Match a character in the set [Ee]. ○ Match a sequence of expressions. ○ Match a character in the set [Ii]. ○ Match a character in the set [Ss]. ○ Match a character in the set [Ii]. ○ Match a character in the set [Nn]. ○ Match a character in the set [Dd]. ○ Match a character in the set [Ee]. ○ Match a character in the set [Xx]. ○ Match a sequence of expressions. ○ Match a character in the set [Ll]. ○ Match a character in the set [Ii]. ○ Match a character in the set [Nn]. ○ Match a character in the set [Kk\u212A]. ○ Match a sequence of expressions. ○ Match a character in the set [Mm]. ○ Match a character in the set [Ee]. ○ Match a character in the set [Tt]. ○ Match a character in the set [Aa]. ○ Match a whitespace character atomically any number of times. ○ Match '>'.

MatchXhtmlMinimizedEndChars()

[GeneratedRegex("\\s*/>")]
public static Regex MatchXhtmlMinimizedEndChars()

Returns

Regex

Remarks

Pattern:

\\s*/>

Explanation:<br />
<pre><code class="lang-csharp">○ Match a whitespace character atomically any number of times.

○ Match the string "/>".

MatchXhtmlSelfClosingTags()

[GeneratedRegex("<\\s*(br|hr|img|link|meta)([^>]*)(>)", RegexOptions.IgnoreCase)]
public static Regex MatchXhtmlSelfClosingTags()

Returns

Regex

Remarks

Pattern:

<\\s*(br|hr|img|link|meta)([^>]*)(>)

Options:<br />
<pre><code class="lang-csharp">RegexOptions.IgnoreCase</code></pre><br />
Explanation:<br />
<pre><code class="lang-csharp">○ Match '<'.

○ Match a whitespace character atomically any number of times. ○ 1st capture group. ○ Match with 5 alternative expressions. ○ Match a sequence of expressions. ○ Match a character in the set [Bb]. ○ Match a character in the set [Rr]. ○ Match a sequence of expressions. ○ Match a character in the set [Hh]. ○ Match a character in the set [Rr]. ○ Match a sequence of expressions. ○ Match a character in the set [Ii]. ○ Match a character in the set [Mm]. ○ Match a character in the set [Gg]. ○ Match a sequence of expressions. ○ Match a character in the set [Ll]. ○ Match a character in the set [Ii]. ○ Match a character in the set [Nn]. ○ Match a character in the set [Kk\u212A]. ○ Match a sequence of expressions. ○ Match a character in the set [Mm]. ○ Match a character in the set [Ee]. ○ Match a character in the set [Tt]. ○ Match a character in the set [Aa]. ○ 2nd capture group. ○ Match a character other than '>' atomically any number of times. ○ 3rd capture group. ○ Match '>'.

MatchXmlNamespaceAttributes()

[GeneratedRegex("\\s*xmlns\\s*=\\s*\"[^\"]+\"\\s*")]
public static Regex MatchXmlNamespaceAttributes()

Returns

Regex

Remarks

Pattern:

\\s*xmlns\\s*=\\s*"[^"]+"\\s*

Explanation:<br />
<pre><code class="lang-csharp">○ Match a whitespace character atomically any number of times.

○ Match the string "xmlns". ○ Match a whitespace character atomically any number of times. ○ Match '='. ○ Match a whitespace character atomically any number of times. ○ Match '"'. ○ Match a character other than '"' atomically at least once. ○ Match '"'. ○ Match a whitespace character atomically any number of times.

PublicDocType(string?, string?, string?)

Emits a public DOCTYPE tag.

public static string PublicDocType(string? rootElement = "html", string? publicIdentifier = "-//W3C//DTD XHTML 1.0 Transitional//EN", string? resourceReference = "http://www.w3.org/TR/xhtml1/DTD/xhtml1-transitional.dtd")

Parameters

rootElement string

The root element of the DTD.

publicIdentifier string

The public identifier of the DTD.

resourceReference string

The link to reference material of the DTD.

Returns

string

A public DOCTYPE tag.