I managed to chat with one of my hero's today, Larry Wall, Godfather of Perl. He was speaking at the Open Source Developer Conference here in Sydney about Perl 6.
Jody Garnett kindly captured the moment on film.
With so many organizations publishing geospatial datasets using standards based web services, a raft of new opportunities for large scale data analysis are presenting themselves. The challenge now is integrating the datasets which use different terms and attributes to describe the same data. For example, "water quality" (good, medium, bad) in one database might equate to "pollution level" (1,2,3,4,5) in another. Communities, like the hydrology community backing the Australian Water Data Infrastructure (AWDIP), are solving these data integration issues by defining a community schema for their domain, then ensuring all agencies publish data using the community schema.
Community Schemas are used to describe a rich set of semantics for a domain, using basic building blocks provided by Geography Markup Language (GML). This allows communities to define schemas appropriate for their data to be used for data transfer within their community. The schemas can then be referenced to ensure consistent structure and taxonomy between related datasets, improving the communities’ ability to share data. For example, a hydrology schema may define a class called Water Use with acceptable terms defined as irrigation, domestic, and industrial. A dataset published with an invalid Water Use of farming will not validate and the user will know to correct the mistake.
Publishing data through a standardised community schema means:
A range of applications and data analysis projects are developed because extensive, quality data becomes available and cost effective
Clear data definitions reduce data misinterpretation
Defining a community schema for a domain is non-trivial as it requires participating parties to create and agree upon a data architecture, vocabularies and an interchange protocol. Luckily, the first community schema projects have left a trail of reusable building blocks and processes that can be used by future efforts. For instance, specifications for Observation and Measurement (O&M) were developed as part of the Sensor Web Enablement (SWE) specification and have since been used as a component in Geoscience Markup Language (GeoSciML), Water Markup Language (WaterML) and others. Other building blocks include Geography Markup Language (GML), SensorML, CityGML and the ANZLIC profile of ISO19115 for Metadata.
The other critical component in the development of a community schema is buy-in, governance and testing from the user community. This is often an international effort. GeoSciML, a schema for geology, has participants from BGS (United Kingdom), BRGM (France), CSIRO (Australia), GA (Australia), GSC (Canada), GSV (Australia), APAT (Italy), JGS (Japan), SGU (Sweden) and USGS (USA) and the OGC (International).
Communities need to adopt a governance structure to resolve the inevitable disagreements over technical details. The GeoSciML community, who started in 2003 and are now onto their third schema iteration, have organised working groups for information model development, computational model development, vocabulary definition, defining use cases, testing the schemas in a formal test bed, then promoting the schema though an outreach working group. A key element in the success of GeoSciML is the fact that custodianship of geologic information is managed by similar agencies in most jurisdictions (Geologic Surveys) and these have a history of collaboration.
Most agencies will first encounter a community schema when they are asked to deploy their datasets using one. Spatial data is collected by numerous agencies, for various purposes, following different collection guidelines. Storage models tend to reflect the original data use and are rarely designed for data exchange. When data is published through web services, the schema usually reflects the storage model. This works fine for the original application but is an integration nightmare when trying to share data between agencies. Changing the storage format usually isn’t desirable if it breaks legacy applications or introduces sub-optimal performance. Hence, it is necessary to differentiate between the storage and exchange formats. Storage format can be defined by the custodian who generates and maintains the data. The challenge is to map the storage model to a community schema. Again, prior projects have built a suite of tools to help out.
Funding from Australia’s National Collaborative Research Infrastructure Strategy (NCRIS), CSIRO and DPI Victoria has added community schema support into GeoServer, an open source WFS and WMS server. Deegree, another open source WFS/WMS is also being investigated. The open source FullMoon supports transforming UML data models into various GML application schemas and is being managed by CSIRO. Duckhawk, an open source WFS & WMS robustness and validation testing tool was developed for the Australia Water Data Infrastructure Project (AWDIP) for testing WaterML.
There is a common theme developing around community schemas; many of the tools being developed are open source. Spatial Data Infrastructures (SDI) increase in value as more agencies contribute data to them. Also, users who benefit most from the infrastructure are often not the agencies that collect or manage the data. This results in large organisations developing SDI’s to aggregate and serve data from numerous smaller, more specialised agencies with different priorities, budgets and timelines. In order to encourage these smaller agencies to manage and publish their data using community schemas, SDI sponsoring organisations like NCRIS are developing Open Source tools in order to reduce the financial barriers faced by small agencies in getting their data online.
Australia, like the rest of the world, has a huge variety of data sets all fulfilling their own purpose and while data integration is non-trivial, we have the knowledge, tools and processes to integrate disparate datasets. This will enable more powerful analysis and new business opportunities for all participating parties.
Credits:
I'd like to thank Stefan Hansen, Software Developer at LISAsoft who was technical lead on the Duckhawk WFS conformance and performance testing framework who helped research this blog. Also Rob Atkinson and Simon Cox who provided a lot of background on CSIRO's involvement with Community Schemas.
A version of this article will be published by Position Magazine in their December 2008 edition.
FOSS4G-2008-LiveDVD-Alpha4.isofrom http://download.osgeo.org/livedvd/ .
Applications:
| GeoServer 1.6.4 | GDAL 1.5.0 | GRASS 6.3.0 |
| gvSIG 1.1.1 | Java 1.6.0 | PgAdmin 1.6.3 |
| PostGIS 1.3.1 | PostgreSQL 8.3.1 | Proj 4.5.0 |
| QGIS 0.11.0 | R 2.4.1 |
Windows Installers:
| FWTools 2.2.1 | gvSIG 1.1.1 |
| ms4w 2.2.8 | QGIS 0.11.0 |
| WinGRASS 6.3.0 |
Data:
| North Carolina GRASS Dataset | QGIS Alaskan Dataset | Spearfish |
| Tasmania Dataset | US States | NYC Dataset |
We, the Mapbuilder Project Steering Committee, have agreed that the time has come for the Community Mapbuilder project to gracefully retire. We will release a final, stable 1.5 version of the software, and afterwards there are no planned enhancements to Mapbuilder. The web pages and code will be kept alive, a few bugs might be fixed and we will likely continue answering user queries, but we expect Mapbuilder will gradually fade away into history.
Mapbuilder is a stable, feature rich, standards compliant, fast, webmapping framework with a strong developer community. Why has it come to the end of its life?
The browser based webmapping space has become crowded and other webmapping clients have increased in functionality and attractiveness to users. In particular, Openlayers is simpler to use, has attracted an incredibly strong developer community, has good quality control and development processes, and has developed most of the webmapping functionality previously only offered by Mapbuilder. Basically Openlayers is attracting the majority of the users and developers that previously would have used Mapbuilder. One day someone will write a compelling paper on the history of the two similar projects and analyse the key differences and decision points which led to one project out shining the other.
Well, maybe we feel a twinge of loss for the Mapbuilder project we started years ago, but in the bigger picture, we see the retiring of Mapbuilder as a good thing. It will allow the greater web mapping community to consolidate and rally around the remaining webmapping tools – in particular, around Openlayers.
There has been significant collaboration between the Mapbuilder and Openlayers communities over the last couple of years. Mapbuilder has incorporated Openlayers as its rendering engine and features have been shared between projects. In many cases, developers from both projects worked on the same codebase (in Openlayers), then ported up to Mapbuilder. This was a deliberate move toward the merging of the two developer communities and most of the Mapbuilder Project Steering Committee have contributed to the Openlayers codebase.
So in essence, by changing our allegiance from Mapbuilder to Openlayers we take with us some of our code, we replace some features with equivalent Openlayers features, we take our community with us, and we gain an existing, robust and welcoming community. We loose a little and gain a lot more.
Users have a few options. You already own the source code, so you are welcome to continue maintaining and extending the Mapbuilder code for as long as you like. At some point, users will likely want to upgrade, and at that point we suggest considering Openlayers for your application. It now provides the majority of the functionality that was previously only offered by Mapbuilder.
Loosing a graduated project might seem like an embarrassment for OSGeo, however, I'd argue it is a strength. It shows two projects growing together under the OSGeo umbrella and eventually merging into a stronger, more focused community.
Even so, it raises a dilemma with regards to what should be done with a retired project. Some key OSGeo criteria like “Community Backing” and “Best of Breed Software” will gradually be lost, so OSGeo should stop recommending Mapbuilder. But, OSGeo shouldn’t erase Mapbuilder's history with OSGeo as Mapbuilder has documented valuable lessons learned during the graduation process.
We suggest OSGeo create a new “retired” project category.
We, the Project Steering Committee, have derived a huge amount of pleasure building Mapbuilder and working with the Mapbuilder Community. For many of us, Mapbuilder has been a launching pad into a full-filling Open Source and/or Geospatial careers. We'd like to thank all the users, developers and supporters of Mapbuilder we have met along the way.
The Mapbuilder Project Steering Committee, (in order of appearance):
Cameron Shorter
Mike Adair
Patrice Cappelaere
Steven M. Ottens
Matt Diez
Olivier Terral
Andreas Hocevar
Gertjan van Oosten
Linda Derezinski
Misapplication of “value for money” requirements when purchasing software results in poor value for money - Government purchasing policies for software tend to support the creation of monopolies.
Government purchasing has effects on the price paid by citizens for the product purchased. In some cases purchasing produces volume which permits scale discounts and therefore a net benefit to citizens who also purchase the product. However, in the case of lock in software Government purchasing can create a monopoly in the software which leads to increased costs for citizen purchasers and a net detriment for society as a whole. It is not appropriate for value for money policies to be assessed on a per acquisition basis when software is being acquired. Doing so will almost certainly create net costs for the community when considered in the aggregate.
Read the rest of Brendan Scott's article here: The Tragedy of the Anti-Commons.
After deciding to extend Open Source Software, agencies are faced with a relatively new business model, open source sponsorship. Agencies need to align purchasing policies, based upon deliverables and milestones, with Open Source community development.
Under a proprietary business model, a company builds and markets a product. Multiple customer sales cover the cost of development, supporting infrastructure, marketing, support, future enhancements and hopefully include a profit. While Open Source business models incur the same costs as the proprietary models they generally distribute the costs to the end users differently, charging for the implementation of specific functional or usability improvements.
Initial investment in communities, infrastructure, and marketing for an Open Source project is often the most effective way to ensure a long term return on investment as these areas are commonly neglected in favour of feature enhancements. Proper promotion and infrastructure support, instead of a sole focus on missing features, will encourage project growth and ultimately lead to open source Nirvana: hundreds of developers building your application using someone else’s budget.
Management should justify investment in a project as Opportunity Management.
Opportunity Management is the inverse of Risk Management. With risk management you quantify what can go wrong then identify mitigation strategies to avoid or reduce the impact of the risks. With opportunity management you list potential windfalls and deploy strategies to enable and benefit from the windfalls. Table 1 shows an example opportunity management matrix.
Table 1: Opportunity Management
| Opportunity | Enabler |
| Use data from external agencies.
| Agencies are given access to open source tools to reduce their barrier to sharing data. Use Open Standards for tools to facilitate communication. Use Open Standards for data schemas so data can be integrated. |
| External Agencies extend our toolset. | Use and share our tools as Open Source Software so that others can use and extend them. Support the Open Source development processes to reduce the barrier of entry to potential development sponsors. |
There are a number of key elements that a potential sponsor should consider when evaluating an open source project in order to ensure maximum return on investment. These include:
Solves a specific need effectively.
Has an active, diverse and inclusive community.
Enjoys support from multiple sponsors.
Established development processes including:
Issue tracking
Communication channels like email lists and IRC
Quality control
Clear and comprehensive documentation and marketing material.
The OpenLayers project is a good example of a commercial entity driving the creation of a thriving open source project. OpenLayers is an open source, browser based web-mapping client which provides a front end to various proprietary and open data sources like Google and Yahoo Maps, WMS and WFS. In three years OpenLayers has grown from nothing to be the dominant open web-mapping client, attracting the majority of the users and developers in this space.
OpenLayers was initially sponsored by MetaCarta who needed a browser based application to support their mapping services. Rather than focusing on features, MetaCarta focused much of their investment on infrastructure and community support. In particular their effort was spent answering developer and user questions on email and IRC, monitoring the quality of code contributions, and setting up automated testing. Many of MetaCartas engineers have developed a personal interest in OpenLayers which MetaCarta encourages by allowing the engineers to spend some work time on the project.
Today, OpenLayers has an incredibly active developer community requiring minimal support from MetaCarta and have provided functionality significantly greater than MetaCarta’s original scope. Key to the success of OpenLayers has been the long running, dedicated community support provided by Chris Schmidt from MetaCarta. GeoServer1, another Open Source project, has recently introduced a similar community liaison role, dedicated to community support and marketing.
The role of Community Liaison has always been key to Open Source and often is filled by volunteer enthusiasts, however commercial deployments of Open Source creates a workload volunteers can’t maintain and hence industry hires these volunteers instead. Ensuring that the community is supported in this fashion promotes the uptake of the project, increases the user base, which in turn attracts more sponsors and more developers. This leads to the situation where many developers are employed by a variety of sponsors to create new features and improve the performance and stability of the project.
There are a number of tasks and roles that need to be addressed in order to ensure a successful open source project. These are described below.
Community Support
A person or team is required to answer user and developer questions, review submitted code from external developers to ensure quality control and ensure that all submissions meet the project requirements in terms of test coverage and documentation. This is one of the most effective investments in a project.
General Project Processes
All projects should invest in tools and processes such as automated build systems, issue trackers, concurrent versioning systems as well as ensuring that releases are performed smoothly and regularly.
Documentation
Good, current design and implementation documentation lowers the learning curve for developers supporting and extending software and greatly increases productivity. Good user documentation engenders confidence in project reviewers which in turn will lead to greater adoption.
Marketing
While Open Source benefits significantly from community generated promotion, it is enhanced by prudent investment in web pages and presentations for targeted conferences.
Commercial Support
One of the main reasons given for avoiding Open Source is not being able to call someone to fix problems. Offering commercial support for a project you use will go far in encouraging adoption by other organizations.
Integrate and bundle with related software
Microsoft Office has been especially successful because it integrates a suite of related products and bundles them all together in one easy install. Open Source products improve their attractiveness in the same way.
Open Standards
Due to the release-early/release-often approach of most open source projects, they are often leveraged to develop, test and extend open standards. This makes open source projects among the earliest adopters of emerging standards, encourages the uptake of open standards and makes the projects attractive to those interested in sharing data between agencies.
Project Management
Just like proprietary software, a sponsor’s software development should be managed using standard software development processes. This includes estimation; planning resources, work activities, schedules, budgets, deliverables; monitoring schedule, quality, risk, issues, contractors, configuration management.
Measurement
Measurement is a key tool used during proprietary Project Management, as good metrics enable good management decisions. Good measures highlight whether specific business goals are being met and enable management to alter their strategy early if issues arise.
Metrics are under-utilized in many open source projects as developers usually drive their own agendas, are self motivated, and spend less time on Project Management. However metrics based decision making can be equally effective for Open Source projects especially for sponsors who will need to answer to commercial milestones and targets.
Standard software development metrics should be complemented by measures to monitor the health of an Open Source community. The Community MapBuilder project tracks many of these metrics. There are now a number of dedicated tools which automate many of the common software metrics.
If your organisation is considering a long term use of an Open Source product, it will likely be smart to invest in the project’s infrastructure. Audit the project against the checklists above, cost the areas needing improvement, set up an opportunity register and determine whether prudent investment will be of value.
Hopefully supporting Open Source will prove to be an effective solution. Then you will discover one more thing,
“Writing Open Source is very good for the soul.”

An Open Source robustness and conformance testing framework will be developed as part of this project. These testing tools will help system integrators and tool developers build quality Spatial Data Infrastructure which is robust and scalable.
The WFSs deployed in AWDIP use the Community Schema functionality which was recently introduced to the stable 1.6 release of Geoserver.
The GML returned by Web Feature Services (WFS) includes semantically defined geometric data describing points, lines polygons as well as other information, like Water Level or Water Use.
Community Schemas define specific fields and acceptable values for the “Other Information” inside GML. Eg: Within a Community Schema we can define Water Use as: agriculture, irrigation, domestic, … but not farming. Using Community Schemas facilitates sharing data between data stores, and enables powerful business applications like the data analysis provided by AWDIP.
More information at on Community Schemas is available at: https://www.seegrid.csiro.au/twiki/bin/view/Infosrvices/InformationViewpoint
Stress, reliability, load, performance, error handling and conformance testing will all be supported. Black Box testing will be used for most of tests, interfacing with the WFS through the Web Service interface.
The tests will be configured using JUnit, allowing it to be incorporated into existing testing and continuous integration frameworks. Project specific tests, as required by AWDIP, will be developed as separate modules.
White box testing will be done for specific AWDIP installs to test internal bottlenecks and database responsiveness.
The Australian Bureau of Rural Sciences is responsible for the Australian Water Data Infrastructure Project (AWDIP), and have contracted LISAsoft to build a deployable testing framework for Geoserver deployments.
The testing framework will build upon prior Geoserver/Mapserver testing done by TOPP and others. In particular, we will use Andrea Aime's blue print for performance tests which integrate into Geoserver's continuous testing framework.
Rob Atkinson (CSIRO Australia) has been leading the effort to support data standards in Geoserver and has been involved in the design and deployment of AWDIP WFS nodes, as well as ongoing support for other international data standards such as GeoSciML.
Duckhawk is the Open Source, WFS testing project. Interested parties are invited to join our email list and monitor or participate in the project.
Standards and tools for reliable data synchronization in Spatial Data Infrastructure and field based data collection.
This article describes the issues and technical solutions associated with Federated Geo-synchronization.
Lisasoft aims to contribute to these solutions as part of the OGC’s Open Web Services Testbed 5.2.
To strengthen and refine our requirements, we are looking for Agencies which would benefit from solutions identified here. Please leave a comment, or contact me if you are interested.
As spatial databases become distributed and collaboratively maintained, traditional database transaction models ineffectively handle modern scenarios.
Figure 1 Synchronising databases in a Spatial Data Infrastructure
Users require current data from remote agencies. Data may be stored on a slow or unreliable server or behind an unreliable internet connection.
Updates may come from remote field workers, trusted external organizations, or general internet users. Identities must be confirmed, updates validated and applied, or rolled back to a previous version.
Figure 2 Caching WFS-T in field, local and remote networks
This project will:
Make a Spatial Data Infrastructure (SDI) fast and robust by caching remote WFSs locally.
Provide WFS Synchronization to allow real time data updates between agencies.
Ensure interoperability between agencies and applications by proposing required extensions to Open Standards.
Ensure wide adoption by providing all components free as Open Source Software.
Provide desktop and mobile, field based data collection tools.
A Transactional Web Feature Service (WFS-T) provides an OGC standards compliant web interface for downloading and updating vector features over the internet. To date the WFS-T standard doesn’t address version history.
Versioned WFS-T enables users to roll back to previous versions, track update history, check differences between updates. A versioned WFS is required to support a cached WFS.
Geoserver developers have developed a Versioned WFS-T by extending the WFS-T specification to include standard version attributes. As at May 2007, the Geotools version code is complete, but still in alpha state. It requires configuration web pages to ease operator use, packaging into a release and real world testing.
The extensions to the WFS-T specification need to go through the OGC standards process.
Security involves: authentication (to verify who a user is) and authorization (to specify what a user can view or update).
Geoserver has prototype authorization and authentication code. Access is provided to the level of WFS. Granular access to a layer or a specific feature is not supported. The code still requires refinement, a user interface and integration with the baseline.
Clients like Udig, Mapbuilder and OpenLayers require security logic. Some of this will be addressed during the Canadian Geographic Data Infrastructure Interoperability Pilot, due to complete October 2007.
A Cached WFS mirrors a remote WFS locally. A Cached WFS-T also caches WFS-T updates when disconnected from the remote WFS.
A Cached WFS-T is used when:
The remote WFS uptime is not guaranteed.
The remote WFS connection is unreliable or unable to handle traffic required.
Cached WFS-T depends upon the Version WFS-T protocol.
An alpha version of Cached WFS (read only) is implemented by Geoserver. A friendly user interface is required to bring this to COTS quality.
Minimal development is required to implement Cached WFS-T (with writes) which has a basic conflict management interface.
Business rules for managing updates and conflicts from disconnected clients will be addressed in a second development phase.
Figure 3 JGrass desktop mapping application
Desktop Mapping offers powerful data manipulation and analysis. There are a number of clients available both proprietary and open source with varying levels of functionality and standards compliance.
A prime candidate for an Open Source Desktop Mapper is UDig and JGrass which are combining forces to provide:
User friendly mapping interface.
Extensive map analysis tools from Grass.
Access to numerous mapping format and data sources from Geotools.
Extensible architecture from Eclipse
Development is required to include:
Embedded cached WFS-T from Geoserver (developed but requires integration and testing)
Embedded data store using H2 (in development)
More work is required to add:
Business logic, views and reports to manage collaborative editing and information from a versioned WFS-T.
Field operators need to create or update geographic data while in the field.
A typical use case involves:
Download geographic data while in the office
Disconnect from the network
Modify, create and delete features and datasets. Interface with a GPS to collect feature information.
Synchronize changes with local or remote data-stores via Standards compliant WFS-T protocol.
Figure 4 Ultra Mobile PC with slide down keyboard
A Tablet or Ruggedized PC provides the same operating environment as a desktop PC. So the Desktop Mapper will port directly to the Tablet.
Integration with a GPS is the only extra development required for the Mobile Tablet.
Figure 5 Mapping on a PDA
PDAs are often used for field work because they are cheaper and smaller than laptops. Along with smaller size they are also less powerful and have less storage capacity.
Cut down versions of Windows (Windows CE) and Java (J2ME) run on most PDAs.
Investigation is required to determine effort required to port the Desktop Mapper to the PDA and whether alternative development would be more effective.
Figure 6 Mapbuilder, a browser map editor/viewer
Browser editors efficiently enable data collection from the public or remote workers.
Browser clients can also publish public map data.
Openlayers and Mapbuilder are working together to produce Open Source, Open Standards Browser Based mapping client. WFS-T editing is supported but needs to include business logic associated with user authentication and access rights.
Multiple agencies tend to run multiple technical solutions. This is fine so long as they interoperate through Open Standards.
The Versioned WFS-T protocol will be presented to the OGC to be formalized as an Open Standard.
Free tools reduce entry costs to a Spatial Data Infrastructure which will maximize participation.
Open Source software already provides the majority of the functionality required by this project which means tools can be built for minimal cost.
Essential Deliverables are required to meet immediate customer needs. These phases are low risk as the functionality already exists in tested or prototype code.
Phase 1: Mirror remote WFS locally
Cached WFS (view only). Builds upon Geoserver/PostGIS.
Standard UDig for WFS viewing
Phase 2: Update remote WFS-T from remote or disconnected client
Cached WFS-T (read/write). Builds upon Geoserver/PostGIS. Add simple update business rules.
Standard UDig for WFS-T editing
Phase 3: Security - Role based editing and views
Security added to Geoserver
Role based options available in UDig
Optional Deliverables are nice to have and involve further development with associated risk.
Phase 4: Universal Client for easy install
UDig with embedded database (H2) and Cached WFS (Geoserver)
Phase 5: Universal Client on PDA
Port universal client to PDA
| OWS 5.1 | |
| RFQ (5.1) Issued | May 11, 2007 |
| OWS 5.2 | |
| Revised RFQ issued | July 9, 2007 |
| Questions Due & Bidders’ Conference | July 16, 2007 (TBR) |
| Clarifications Posted | July 23, 2007 (TBR) |
| RFQ Responses Due | August 3, 2007 |
| Kickoff Meeting | week of September 10, 2007 |
| Interim Milestone | week of November 12, 2007 |
| Demonstration Milestone | week of January 7, 2008 |
| Final Delivery | February 18 – February 22, 2008 |
| Commercialize product, provide support, consulting and customized solutions. | March 2008 – 2009. |
Using free Open Source allows Systems Integrators to increase services or reduce price.
... I was busy writing and delivering a new course for PennState on Open Web Mapping. Finally its all over and its time to give back to the community first there is a page of student projects the majority are a built with GeoServer and MapBuilder at http://webmapping.mgis.psu.edu/mapbuilder/demo/index2.html. PennState has also generously agreed to give away the course ware under a CCSA license so you can all see what I've been saying about your projects at https://courseware.e-education.psu.edu/courses/geog585/content/home.html. If any one would like to take the two case study lessons and roll them in to tutorials you're welcome.
In general all the students were very happy about the quality and ease of use of the open source tools they used, mostly they wanted more MapBuilder documentation and more projections.
Background