Replica Management in Data Grids

Providing fast, reliable and transparent access to data to all users within a community is one of the the most crucial functions of data management in a Grid environment. User communities are typically large and highly geographically distributed. The volume of data that they wish to access is of the order of petabytes and may also be distributed. It is infeasible for all users to access a single instance of all data. One solution is that of data replication. Identical replicas of data are generated and stored at various globally distributed sites. Replication can reduce data access latency and increase the performance and robustness of distributed applications. The existence of multiple instances of data however introduces additional issues. Replicas must be kept consistent, they must be locatable, and their lifetime must be managed. These and other issues necessitate a high level system for replica management in Data Grids. This paper describes the architecture and design of a Replica Management System, called Reptor, within the context of the EU Data Grid project. A prototype implementation is currently under development. 1

Replica Management in Data Grids | Litlas