| The
ability to use some of the built in methods of the Assembly class, such
as GetManifestResourceStream, makes it easy for developers
to embed resources of almost any type into an assembly and access these
at runtime to be used in the functionality of the class itself.
You can
embed almost any type of resource - images, Javascript files, text representations
of XmlDocuments, and even delimited text files that have been exported
from a database. But what about the idea of using Data Compression on
some of these resources before they are embedded in your
assembly? Based on some of the work I've done on compression, text files
can often be compressed by as much as 90 percent. If you had the code
to compress and decompress these resources, and the decompression code
could be included in your assembly, then you would have the ingredients
to access and decompress your in-assembly resources at runtime and since
they would already be memory-resident at the time of decompression, you
could potentially save a considerable amount of time.
In the
case of an ASP.NET application, the one-time loading of the assembly could
be stored in Application state, and would then be accessible to any ASP.NET
page running in the application with no loading delays. And in the case
of other applications, the rehydrated resource could be stored in the
appDomain cache instead, basically accomplishing the same effect.
As it
turns out, Mike Kruger's SharpZipLib C# codec provides the source code
we need, and we would only need to include those classes necessary to
decompress an embedded resource in order to meet our needs, so the impact
of the decompression code on the size of our final assembly is minimal.
As a
"Proof of Concept" exercise, I took an export of a database
we have of famous quotations; said export consisting of 8,627 records
each containing an Author, Author Info, Subject and a Quotation column,
and put this into a delimited CSV text file which takes up 1186 Kb on
the hard drive. After applying standard compression with a maximum compression
factor from SharpZipLib, the resultant resource is only 423K - a 64% space
savings. Here are a couple of the uncompressed records so that you can
see what I started with:
Meister
Eckhart|1260-1326 AD, German Mystic|Compassion|You may call God love,
you may call God goodness. But the best name for God is compassion.
Dante (Alighieri)|1265-1321, Italian Philosopher, Poet|Conscience|O conscience,
upright and stainless, how bitter a sting to thee is a little fault!
Along
with all the code to decompress this embedded resource into a DataTable
and provide two separate DataTable search methods FindQuotesByAuthor,
and FindQuotesBySubject, each using the Select filter method on the rehydrated
DataTable in memory, the final QuotationControl assembly is just 472K
- well within the range of acceptable DLL sizes to be loaded into memory
one time upon Application instantiation.
The
Select method is very fast, even though the DataTable is not indexed,
because the records are all in-memory and so there is no disk I/O overhead
as with a database. I've left out primary keys and indexing here in order
to keep the focus on the basic premise, which is that of being able to
decompress an embedded resource at runtime.
However,
though ADO.NET does not expose indexes, it does create and use them both
for Primary Keys and for DataViews. You can take advantage of the indexes
created for Primary Keys or DataViews as follows:
If
you are searching on a Primary Key field, or if you can make the field(s)
you are searching on a Primary Key, then instead of DataTable.Select(),
you can use DataTable.Rows.Find(). Find() will look up
a primary key value using the built index.
If you are performing repeated searches
on a particular column or set of columns, you can create a DataView sorted
on those columns, and then use DataView.FindRows() to
search for a value within the columns. DataViews create an index for the
sorted columns, and FindRows() uses this. Also, when using a DataView,
it is more efficient to use the constructor that accepts the table, filter,
sort, and rowversions all at once. Each time you set the filter, sort,
or rowversion properties the index is recreated, so creating a DataView
and setting these properties individually actually creates the index several
times. An example is shown immediately below:
private
void MakeDataView(DataSet ds)
{
DataView dv = new DataView(ds.Tables["Suppliers"], "Country
= 'UK'", "CompanyName", DataViewRowState.CurrentRows);
dv.AllowEdit = true;
dv.AllowNew = true;
dv.AllowDelete = true;
}
I've included two solutions in the downloadable
source code for this article:
The
first solution is called NZipLibForm and is just an updated test harness
for SharpZipLib that allows you to choose Compress, get a file dialog,
find the resource you want compressed, and have it automatically saved
in the same folder with the extension ".dat" appended to it.
By leaving the "Compress to/from file" checkbox unchecked, you
can also paste a resource into the upper text box and see the compressed
result in the lower textbox.
The second
solution is the QuotationControl (its not really a control, just a class
library, but that's the way I started it out so I've been reluctant to
rename it). This is the real meat of the madness, and basically looks
like this:
using System;
using System.Web.UI;
using System.Reflection;
using System.IO;
using System.Data;
using System.Diagnostics;
using ICSharpCode.SharpZipLib;
using System.Collections;
namespace QuotationControl
{
public class QuotationLookup
{
private string theQuotes=String.Empty;
public DataTable QuotesTable = new DataTable();
private DataTable foundQuotes=new DataTable();
public DataTable FoundQuotes
{
get
{
return foundQuotes;
}
}
public void FindQuotesByAuthor
(string Author)
{
foundQuotes.Clear();
string strExpr;
strExpr = "Author LIKE '%" +Author.Trim() + "%'";
// Use the Select method to find all rows matching the filter.
try
{
DataRow[] foundRows =
QuotesTable.Select(strExpr);
for(int j =0;j<foundRows.Length;j++)
{
object[] theRow={foundRows[j].ItemArray[0].ToString(),
foundRows[j].ItemArray[1].ToString(),
foundRows[j].ItemArray[2].ToString(),
foundRows[j].ItemArray[3].ToString()};
foundQuotes.Rows.Add(theRow);
}
foundQuotes.AcceptChanges();
}
catch(Exception ex)
{throw new Exception(ex.Message);
}
}
public void FindQuotesBySubject (string Subject)
{
foundQuotes.Clear();
string strExpr;
strExpr = "Subject LIKE '%" +Subject.Trim() + "%'";
try
{
DataRow[] foundRows =
QuotesTable.Select(strExpr);
for(int j =0;j<foundRows.Length;j++)
{
object[] theRow={foundRows[j].ItemArray[0].ToString(),
foundRows[j].ItemArray[1].ToString(),
foundRows[j].ItemArray[2].ToString(),
foundRows[j].ItemArray[3].ToString()};
foundQuotes.Rows.Add(theRow);
}
foundQuotes.AcceptChanges();
}
catch(Exception ex)
{
throw new Exception(ex.Message);
}
}
public QuotationLookup()
{
foundQuotes.Columns.Add("Author");
foundQuotes.Columns.Add("AuthorInfo");
foundQuotes.Columns.Add("Subject");
foundQuotes.Columns.Add("Quotation");
theQuotes = GetDecompressedResourceString("QuotationControl.quotations.dat");
QuotesTable.Columns.Add( "Author", typeof(string) );
QuotesTable.Columns.Add( "AuthorInfo", typeof(string)
);
QuotesTable.Columns.Add( "Subject", typeof(string) );
QuotesTable.Columns.Add( "Quotation", typeof(string) );
theQuotes=theQuotes.Replace("\"","");
string[] whizzies = theQuotes.Split(new Char[] {'\n'});
for (int i=0;i<whizzies.Length;i++)
{
object[] theRow=whizzies[i].Split(new Char[] {'|'});
QuotesTable.Rows.Add( theRow);
}
QuotesTable.AcceptChanges();
}
private string GetDecompressedResourceString(string
resource)
{
try
{
Assembly asm = Assembly.GetExecutingAssembly();
Debug.WriteLine(asm.FullName);
Stream stm=asm.GetManifestResourceStream(resource);
Debug.WriteLine(stm.Length.ToString());
BinaryReader br = new BinaryReader(stm);
long siz = stm.Length;
byte[] bytInput =null;
bytInput=br.ReadBytes((int)siz);
string theResource=DeCompress(bytInput);
br.Close();
stm.Close();
return theResource;
}
catch(Exception ex)
{
throw new ApplicationException(ex.Message);
}
}
private string DeCompress(byte[]
bytInput)
{
string strResult="";
int totalLength = 0;
byte[] writeData = new byte[4096];
Stream s2 = new ICSharpCode.SharpZipLib.Zip.Compression.Streams.InflaterInputStream(new
MemoryStream(bytInput));
try
{
while (true)
{
int size = s2.Read(writeData, 0, writeData.Length);
if (size > 0)
{
totalLength += size;
strResult+=System.Text.Encoding.UTF8.GetString(writeData, 0,
size);
}
else
{
break;
}
}
s2.Close();
return strResult;
}
catch(Exception e)
{
throw new Exception(e.ToString());
}
}
}
}
|
Note
that I've left out the SharpZipLib code classes in the above snippet,
but basically the DeCompress method encapsulates all the code necessary
to invoke the decompression engine on your passed in resource stream.
Note also that it is necessary to use the BinaryReader
to get your resource out of the assembly in the correct format for decompression
to work!
Once
the resource is decompressed into a string, it is a relatively straightforward
process to use the Split method, once on the newline character to get
all your rows into an array, and then again on the pipe ("|")
character which is what I chose as my column delimiter when exporting
these records to a text file from Sql Server, in order to get the column
items for each row in the DataTable.
You
can apply this overall technique to tabular data such as the above (you
could also construct an entire DataSet in memory complete with indexes,
primary keys, and constraints), it can be applied to a large XmlDocument
which is then rehydrated and stored in - memory, and so on.
If you'd
like to try out the result before downloading the sample solutions, just
hit
this demo page. For an Author search, you might try the search term
"shak" (for Shakespeare) or "Bacon" (for Francis Bacon),
and for a subject search, why not try "Happiness" - a good subject
for all of us! Interestingly, as a philosophical side note, since my search
term is surrounded by "%" wildcards, if you search on "war"
you will also be presented with the quotations on the subjects of "cowardice"
and "rewards". Enjoy.
Download
the Code that accompanies this article
| | Peter Bromberg is a C# MVP, MCP, and .NET consultant who has worked in the banking and financial industry for 20 years. He has architected and developed web - based corporate distributed application solutions since 1995, and focuses exclusively on the .NET Platform. |
|