Class HtmlUtility
Static members for HTML text processing.
public static class HtmlUtility
- Inheritance
-
HtmlUtility
- Inherited Members
Methods
ConvertToHtml(string?)
Returns a string of marked up text compatible with browsers that do not support XHTML (loosely towards HTML 4.x W3C standard).
public static string? ConvertToHtml(string? input)
Parameters
Returns
ConvertToXml(string?)
Attempts to convert HTML to well-formed XML with a conventional set of Regex methods.
public static string? ConvertToXml(string? html)
Parameters
Returns
Remarks
This method is simpler than converting to XHTML.
This member uses the following Regex methods:
- MatchXmlNamespaceAttributes()
- MatchXhtmlSelfClosingTags()
- MatchHtmlStartTags() (for two different evaluations)
- MatchHtmlHrefOrSrcAttributes()
FormatXhtmlElements(string?)
Formats the specified fragment with a conventional set of Regex methods.
public static string? FormatXhtmlElements(string? xmlFragment)
Parameters
Returns
Remarks
This member uses the following Regex method:
GetBooleanAttributes()
Returns all known HTML Boolean attributes
public static string[] GetBooleanAttributes()
Returns
- string[]
Remarks
GetHtmlBooleanAttributesPattern()
Regex pattern getter
public static string GetHtmlBooleanAttributesPattern()
Returns
GetInnerXml(string?, string?, string, byte)
Returns the …inner… fragment of XML from the specified unique element.
public static string? GetInnerXml(string? xmlFragment, string? elementName, string newLine = "\r\n", byte numberOfChars = 4)
Parameters
xmlFragmentstringA well-formed string of XML.
elementNamestringThe local name of the element in the XML string.
newLinestringthe conventional NewLine characters of the lines
numberOfCharsbytethe number of white space characters to remove
Returns
MatchHtmlAttributesWithoutQuotes()
[GeneratedRegex("\\s([^'\\\"\\s]+)(\\s*=\\s*)([^'\\\"\\s]+)\\s", RegexOptions.IgnoreCase)]
public static Regex MatchHtmlAttributesWithoutQuotes()
Returns
Remarks
Pattern:
\\s([^'\\"\\s]+)(\\s*=\\s*)([^'\\"\\s]+)\\sOptions:<br />
<pre><code class="lang-csharp">RegexOptions.IgnoreCase</code></pre><br />
Explanation:<br />
<pre><code class="lang-csharp">○ Match a whitespace character.
○ 1st capture group. ○ Match a character in the set [^"'\s] greedily at least once. ○ 2nd capture group. ○ Match a whitespace character atomically any number of times. ○ Match '='. ○ Match a whitespace character greedily any number of times. ○ 3rd capture group. ○ Match a character in the set [^"'\s] greedily at least once. ○ Match a whitespace character.
MatchHtmlClosingTagCharacters()
[GeneratedRegex("\\<\\s*/")]
public static Regex MatchHtmlClosingTagCharacters()
Returns
Remarks
Pattern:
\\<\\s*/Explanation:<br />
<pre><code class="lang-csharp">○ Match '<'.
○ Match a whitespace character atomically any number of times. ○ Match '/'.
MatchHtmlElementsThatShouldNotBeMinimized()
[GeneratedRegex("<(a|iframe|td|th|script)\\s+([^>]*)(/>)", RegexOptions.IgnoreCase)]
public static Regex MatchHtmlElementsThatShouldNotBeMinimized()
Returns
Remarks
Pattern:
<(a|iframe|td|th|script)\\s+([^>]*)(/>)Options:<br />
<pre><code class="lang-csharp">RegexOptions.IgnoreCase</code></pre><br />
Explanation:<br />
<pre><code class="lang-csharp">○ Match '<'.
○ 1st capture group. ○ Match with 5 alternative expressions. ○ Match a character in the set [Aa]. ○ Match a sequence of expressions. ○ Match a character in the set [Ii]. ○ Match a character in the set [Ff]. ○ Match a character in the set [Rr]. ○ Match a character in the set [Aa]. ○ Match a character in the set [Mm]. ○ Match a character in the set [Ee]. ○ Match a sequence of expressions. ○ Match a character in the set [Tt]. ○ Match a character in the set [Dd]. ○ Match a sequence of expressions. ○ Match a character in the set [Tt]. ○ Match a character in the set [Hh]. ○ Match a sequence of expressions. ○ Match a character in the set [Ss]. ○ Match a character in the set [Cc]. ○ Match a character in the set [Rr]. ○ Match a character in the set [Ii]. ○ Match a character in the set [Pp]. ○ Match a character in the set [Tt]. ○ Match a whitespace character greedily at least once. ○ 2nd capture group. ○ Match a character other than '>' greedily any number of times. ○ 3rd capture group. ○ Match the string "/>".
MatchHtmlHrefOrSrcAttributes()
[GeneratedRegex("(href|src)\\s*=\\s*['\"][^\"]+['\"]", RegexOptions.IgnoreCase)]
public static Regex MatchHtmlHrefOrSrcAttributes()
Returns
Remarks
Pattern:
(href|src)\\s*=\\s*['"][^"]+['"]Options:<br />
<pre><code class="lang-csharp">RegexOptions.IgnoreCase</code></pre><br />
Explanation:<br />
<pre><code class="lang-csharp">○ 1st capture group.
○ Match with 2 alternative expressions.
○ Match a sequence of expressions.
○ Match a character in the set [Hh].
○ Match a character in the set [Rr].
○ Match a character in the set [Ee].
○ Match a character in the set [Ff].
○ Match a sequence of expressions.
○ Match a character in the set [Ss].
○ Match a character in the set [Rr].
○ Match a character in the set [Cc].
○ Match a whitespace character atomically any number of times. ○ Match '='. ○ Match a whitespace character atomically any number of times. ○ Match a character in the set ["']. ○ Match a character other than '"' greedily at least once. ○ Match a character in the set ["'].
MatchHtmlStartTags()
[GeneratedRegex("<[^>\\/]+>")]
public static Regex MatchHtmlStartTags()
Returns
Remarks
Pattern:
<[^>\\/]+>Explanation:<br />
<pre><code class="lang-csharp">○ Match '<'.
○ Match a character in the set [^/>] atomically at least once. ○ Match '>'.
MatchHtmlTagContents()
[GeneratedRegex("<[^/][^>]*>")]
public static Regex MatchHtmlTagContents()
Returns
Remarks
Pattern:
<[^/][^>]*>Explanation:<br />
<pre><code class="lang-csharp">○ Match '<'.
○ Match any character other than '/'. ○ Match a character other than '>' atomically any number of times. ○ Match '>'.
MatchHtmlTagWithAnyAttributes()
[GeneratedRegex("<html [^>]*>", RegexOptions.IgnoreCase)]
public static Regex MatchHtmlTagWithAnyAttributes()
Returns
Remarks
Pattern:
<html [^>]*>Options:<br />
<pre><code class="lang-csharp">RegexOptions.IgnoreCase</code></pre><br />
Explanation:<br />
<pre><code class="lang-csharp">○ Match '<'.
○ Match a character in the set [Hh]. ○ Match a character in the set [Tt]. ○ Match a character in the set [Mm]. ○ Match a character in the set [Ll]. ○ Match ' '. ○ Match a character other than '>' atomically any number of times. ○ Match '>'.
MatchXhtmlAttribute()
[GeneratedRegex("\\s+(.+)\\s*=\\s*\"\\1\"")]
public static Regex MatchXhtmlAttribute()
Returns
Remarks
Pattern:
\\s+(.+)\\s*=\\s*"\\1"Explanation:<br />
<pre><code class="lang-csharp">○ Match a whitespace character greedily at least once.
○ 1st capture group. ○ Match a character other than '\n' greedily at least once. ○ Match a whitespace character atomically any number of times. ○ Match '='. ○ Match a whitespace character atomically any number of times. ○ Match '"'. ○ Match the same text as matched by the 1st capture group. ○ Match '"'.
MatchXhtmlEndTagsThatShouldBeMinimized()
[GeneratedRegex("\\s*</(base|isindex|link|meta)\\s*>", RegexOptions.IgnoreCase)]
public static Regex MatchXhtmlEndTagsThatShouldBeMinimized()
Returns
Remarks
Pattern:
\\s*</(base|isindex|link|meta)\\s*>Options:<br />
<pre><code class="lang-csharp">RegexOptions.IgnoreCase</code></pre><br />
Explanation:<br />
<pre><code class="lang-csharp">○ Match a whitespace character atomically any number of times.
○ Match the string "</". ○ 1st capture group. ○ Match with 4 alternative expressions. ○ Match a sequence of expressions. ○ Match a character in the set [Bb]. ○ Match a character in the set [Aa]. ○ Match a character in the set [Ss]. ○ Match a character in the set [Ee]. ○ Match a sequence of expressions. ○ Match a character in the set [Ii]. ○ Match a character in the set [Ss]. ○ Match a character in the set [Ii]. ○ Match a character in the set [Nn]. ○ Match a character in the set [Dd]. ○ Match a character in the set [Ee]. ○ Match a character in the set [Xx]. ○ Match a sequence of expressions. ○ Match a character in the set [Ll]. ○ Match a character in the set [Ii]. ○ Match a character in the set [Nn]. ○ Match a character in the set [Kk\u212A]. ○ Match a sequence of expressions. ○ Match a character in the set [Mm]. ○ Match a character in the set [Ee]. ○ Match a character in the set [Tt]. ○ Match a character in the set [Aa]. ○ Match a whitespace character atomically any number of times. ○ Match '>'.
MatchXhtmlMinimizedEndChars()
[GeneratedRegex("\\s*/>")]
public static Regex MatchXhtmlMinimizedEndChars()
Returns
Remarks
Pattern:
\\s*/>Explanation:<br />
<pre><code class="lang-csharp">○ Match a whitespace character atomically any number of times.
○ Match the string "/>".
MatchXhtmlSelfClosingTags()
[GeneratedRegex("<\\s*(br|hr|img|link|meta)([^>]*)(>)", RegexOptions.IgnoreCase)]
public static Regex MatchXhtmlSelfClosingTags()
Returns
Remarks
Pattern:
<\\s*(br|hr|img|link|meta)([^>]*)(>)Options:<br />
<pre><code class="lang-csharp">RegexOptions.IgnoreCase</code></pre><br />
Explanation:<br />
<pre><code class="lang-csharp">○ Match '<'.
○ Match a whitespace character atomically any number of times. ○ 1st capture group. ○ Match with 5 alternative expressions. ○ Match a sequence of expressions. ○ Match a character in the set [Bb]. ○ Match a character in the set [Rr]. ○ Match a sequence of expressions. ○ Match a character in the set [Hh]. ○ Match a character in the set [Rr]. ○ Match a sequence of expressions. ○ Match a character in the set [Ii]. ○ Match a character in the set [Mm]. ○ Match a character in the set [Gg]. ○ Match a sequence of expressions. ○ Match a character in the set [Ll]. ○ Match a character in the set [Ii]. ○ Match a character in the set [Nn]. ○ Match a character in the set [Kk\u212A]. ○ Match a sequence of expressions. ○ Match a character in the set [Mm]. ○ Match a character in the set [Ee]. ○ Match a character in the set [Tt]. ○ Match a character in the set [Aa]. ○ 2nd capture group. ○ Match a character other than '>' atomically any number of times. ○ 3rd capture group. ○ Match '>'.
MatchXmlNamespaceAttributes()
[GeneratedRegex("\\s*xmlns\\s*=\\s*\"[^\"]+\"\\s*")]
public static Regex MatchXmlNamespaceAttributes()
Returns
Remarks
Pattern:
\\s*xmlns\\s*=\\s*"[^"]+"\\s*Explanation:<br />
<pre><code class="lang-csharp">○ Match a whitespace character atomically any number of times.
○ Match the string "xmlns". ○ Match a whitespace character atomically any number of times. ○ Match '='. ○ Match a whitespace character atomically any number of times. ○ Match '"'. ○ Match a character other than '"' atomically at least once. ○ Match '"'. ○ Match a whitespace character atomically any number of times.
PublicDocType(string?, string?, string?)
Emits a public DOCTYPE tag.
public static string PublicDocType(string? rootElement = "html", string? publicIdentifier = "-//W3C//DTD XHTML 1.0 Transitional//EN", string? resourceReference = "http://www.w3.org/TR/xhtml1/DTD/xhtml1-transitional.dtd")
Parameters
rootElementstringThe root element of the DTD.
publicIdentifierstringThe public identifier of the DTD.
resourceReferencestringThe link to reference material of the DTD.
Returns
- string
A public
DOCTYPEtag.