hadoop 分布式缓存-阿里云开发者社区

hadoop 分布式缓存

2017-12-04 1093

版权

本文内容由阿里云实名注册用户自发贡献，版权归原作者所有，阿里云开发者社区不拥有其著作权，亦不承担相应法律责任。具体规则请查看《阿里云开发者社区用户服务协议》和《阿里云开发者社区知识产权保护指引》。如果您发现本社区中有涉嫌抄袭的内容，填写侵权投诉表单进行举报，一经查实，本社区将立刻删除涉嫌侵权内容。

简介：

Hadoop 分布式缓存实现目的是在所有的MapReduce调用一个统一的配置文件，首先将缓存文件放置在HDFS中，然后程序在执行的过程中会可以通过设定将文件下载到本地具体设定如下：

public static void main(String[] arge) throws IOException, ClassNotFoundException, InterruptedException{

       Configuration conf=new Configuration();
       conf.set("fs.default.name", "hdfs://192.168.1.45:9000");
       FileSystem fs=FileSystem.get(conf);
       fs.delete(new Path("CASICJNJP/gongda/Test_gd20140104"));

       conf.set("mapred.job.tracker", "192.168.1.45:9001");
       conf.set("mapred.jar", "/home/hadoop/workspace/jar/OBDDataSelectWithImeiTxt.jar");
       Job job=new Job(conf,"myTaxiAnalyze");


        DistributedCache.createSymlink(job.getConfiguration());//
       try {
           DistributedCache.addCacheFile(new URI("/user/hadoop/CASICJNJP/DistributeFiles/imei.txt"), job.getConfiguration());
       } catch (URISyntaxException e1) {
           // TODO Auto-generated catch block
           e1.printStackTrace();
       }
       job.setMapperClass(OBDDataSelectMaper.class);
       job.setReducerClass(OBDDataSelectReducer.class);
       //job.setNumReduceTasks(10);
       //job.setCombinerClass(IntSumReducer.class);
       job.setMapOutputKeyClass(Text.class);
       job.setMapOutputValueClass(Text.class);

       FileInputFormat.addInputPath(job, new Path("/user/hadoop/CASICJNJP/SortedData/20140104"));
       FileOutputFormat.setOutputPath(job, new Path("CASICJNJP/gongda/SelectedData"));

       System.exit(job.waitForCompletion(true)?0:1);

   }

代码中标红的为将HDFS中的/user/hadoop/CASICJNJP/DistributeFiles/imei.txt作为分布式缓存

public class OBDDataSelectMaper extends Mapper<Object, Text, Text, Text> {
   String[] strs;
   String[] ImeiTimes;
   String timei;
   String time;
   private java.util.List<Integer> ImeiList = new java.util.ArrayList<Integer>();

   protected void setup(Context context) throws IOException,
           InterruptedException {

        try {
           Path[] cacheFiles = DistributedCache.getLocalCacheFiles(context
                   .getConfiguration());
           if (cacheFiles != null && cacheFiles.length > 0) {
               String line;
               BufferedReader br = new BufferedReader(new FileReader(
                       cacheFiles[0].toString()));
               try {
                   line = br.readLine();
                   while ((line = br.readLine()) != null) {
                       ImeiList.add(Integer.parseInt(line));
                   }
               } finally {
                   br.close();
               }
           }
       } catch (IOException e) {
           System.err.println("Exception reading DistributedCache: " + e);
       }
   }

   public void map(Object key, Text value, Context context)
           throws IOException, InterruptedException {

       try {
           strs = value.toString().split("\t");
           ImeiTimes = strs[0].split("_");
           timei = ImeiTimes[0];
           if (ImeiList.contains(Integer.parseInt(timei))) {
               context.write(new Text(strs[0]), value);
           }
       } catch (Exception ex) {

       }
   }
}

上述标红代码中在Map的setup函数中加载分布式缓存。

本文转自博客园知识天地的博客，原文链接：hadoop 分布式缓存，如需转载请自行联系原博主。

hadoop 分布式缓存

热门文章

最新文章

相关课程

相关电子书

相关实验场景